VictoriaMetrics

mirror of https://github.com/VictoriaMetrics/VictoriaMetrics.git synced 2024-12-18 14:40:26 +01:00

Author	SHA1	Message	Date
Hui Wang	1908d44cf8	dashboards: fix description about pending datapoints (#7235 ) Some checks failed build / Build (push) Has been cancelled Details CodeQL Go / Analyze (push) Has been cancelled Details CodeQL JS/TS / Analyze (push) Has been cancelled Details main / lint (push) Has been cancelled Details main / test (test-full) (push) Has been cancelled Details main / test (test-full-386) (push) Has been cancelled Details main / test (test-pure) (push) Has been cancelled Details See [our playground](https://play-grafana.victoriametrics.com/d/oS7Bi_0Wz_vm/victoriametrics-cluster-vm?orgId=1&var-ds=P996FABE17B5F6D1E&var-job=All&var-job_insert=All&var-job_select=All&var-job_storage=All&var-instance=All) for reference. (cherry picked from commit `d3f110373c`)	2024-10-11 14:28:23 +02:00
Roman Khavronenko	deb2f87074	deployment: add panel and alerts for displying go scheduler latency (#7078 ) The panel and alerting rule should help to understand whether VM component doesn't have enough CPU resources or gets throttled. The alert is applicable for all VM components. The panel was added to vmalert, vmagent, vmsingle, vm clusert and victorialogs dashes. ------------------- This alerting rule should have help us identify resource shortage for sandbox vmagent - see [this link](https://play.victoriametrics.com/select/accounting/1/6a716b0f-38bc-4856-90ce-448fd713e3fe/prometheus/graph/#/?g0.range_input=23d13h25m25s424ms&g0.end_input=2024-09-23T14%3A11%3A00&g0.relative_time=none&g0.tab=0&g0.expr=histogram_quantile%280.99%2C+sum%28rate%28go_sched_latencies_seconds_bucket%7Bjob%3D%22vmagent-monitoring-vmagent%22%7D%5B5m%5D%29%29+by+%28le%2C+job%2C+instance%29%29+%3E+0.1) for example. We weren't aware of resource shortage, because VM metrics assumed this vmagent had 1vCPU while in fact its limit was 0.2vCPU. Signed-off-by: hagen1778 <roman@victoriametrics.com> (cherry picked from commit `4d0b41e63b`)	2024-09-24 16:58:14 +02:00
zjbztianya	42ad757ac4	dashboards: typo fix (#6920 ) ### Describe Your Changes Correct the spelling error of 'vminsert' in the dashboards. ### Checklist The following checks are mandatory: - [ ] My change adheres [VictoriaMetrics contributing guidelines](https://docs.victoriametrics.com/contributing/). (cherry picked from commit `1b1e61030b`)	2024-09-03 10:49:32 +02:00
hagen1778	89819f2054	dashboards: use `$__interval` variable for offsets and look-behind windows in annotations This should improve precision of `restarts` and `version change` annotations when zooming-in/zooming-out on the dashboards. The change also makes `restarts` dashboard visible on the panels, so user can disable it from displaying if needed. This could be useful when restarts overlap with version change events. Signed-off-by: hagen1778 <roman@victoriametrics.com> (cherry picked from commit `9dd9b4442f`)	2024-05-22 16:40:08 +02:00
hagen1778	0dd3fec2b7	deployment/dashboards: fix `AnnotationQueryRunner` error in Grafana The error appears when executing annotations query against Prometheus backend because the query itself hasn't specified look-behind window (which is allowed in VictoriaMetrics query engine). https://github.com/VictoriaMetrics/VictoriaMetrics/issues/6309 Signed-off-by: hagen1778 <roman@victoriametrics.com> (cherry picked from commit `c746ba154d`)	2024-05-21 16:37:23 +02:00
hagen1778	d87c8757cf	dashboards: add new panel `Concurrent selects` to `vmstorage` row The panel will show how many ongoing select queries are processed by vmstorage and should help to identify resource bottlenecks. See panel description for more details. Signed-off-by: hagen1778 <roman@victoriametrics.com> (cherry picked from commit `d386a68b59`)	2024-04-30 10:30:08 +02:00
hagen1778	0f72ab8ef6	deployment: bump Grafana version to 10.4.2 Signed-off-by: hagen1778 <roman@victoriametrics.com> (cherry picked from commit `9256df17fa`)	2024-04-30 10:30:00 +02:00
hagen1778	d4b56d467f	dashboards: show max number of active merges instead of cumulative The cumulative number of active merges could be red herring as it its value depends on the number of vmstorages. For example, vmstorage could be added or removed and this will affect the panel. Or, each vmstorage could start a merging process (i.e. for downsampling) and visiually it could look like a massive change. Signed-off-by: hagen1778 <roman@victoriametrics.com> (cherry picked from commit `035de57e5e`)	2024-04-24 17:08:15 +02:00
Aliaksandr Valialkin	a21d1fcf57	all: replace old https://docs.victoriametrics.com/Cluster-VictoriaMetrics.html url with the new one - https://docs.victoriametrics.com/cluster-victoriametrics/	2024-04-18 02:56:28 +02:00
Aliaksandr Valialkin	64938732e3	all: replace old https://docs.victoriametrics.com/MetricsQL.html url with the new one - https://docs.victoriametrics.com/metricsql/	2024-04-18 02:15:33 +02:00
hagen1778	764fc566ff	dashboards: add more context to cluster dashboard panels Signed-off-by: hagen1778 <roman@victoriametrics.com>	2024-03-06 13:34:10 +02:00
hagen1778	51745ec5ff	dashboards: update links in various panels * use docs.victoriametrics.com instead of github docs * add links to common terms used in VictoriaMetrics Signed-off-by: hagen1778 <roman@victoriametrics.com>	2024-03-04 17:00:54 +02:00
hagen1778	f4578826b3	dashboards: add legend details to network panels in cluster dash Signed-off-by: hagen1778 <roman@victoriametrics.com> (cherry picked from commit `ecccd2a1cc`)	2024-02-16 15:31:57 +01:00
hagen1778	9b173c2f01	dashboards: follow-up `4369bc1df2` * add more details to changelog * simplify panels description * remove capacity planning recommendation, as it proves it incompetent Signed-off-by: hagen1778 <roman@victoriametrics.com>	2024-02-08 12:55:42 +02:00
Hui Wang	0cd0ddc1c1	deployment/dashboards: fix `Storage full ETA` panels (#5747 ) During background downsampling, rate(vm_deduplicated_samples_total{type="merge"}) could be much bigger than rate(vm_rows_added_to_storage_total) and it could last quite some time, which causes negative values of Storage full ETA and confuses users, see playground. Instead of trying to get more accurate results during downsampling, I think it's ok to ignore vm_deduplicated_samples_total at all, it's more reasonable to see Storage full ETA increase after downsampling. --------- Co-authored-by: Aliaksandr Valialkin <valyala@victoriametrics.com>	2024-02-08 12:54:31 +02:00
hagen1778	bdbab7bed5	dashboards/all: add new panel `CPU spent on GC` It should help identifying cases when too much CPU is spent on garbage collection, and advice users on how this can be addressed. Signed-off-by: hagen1778 <roman@victoriametrics.com>	2024-02-05 11:42:28 +02:00
hagen1778	3dab94a6c1	dashboards: update to grafana/grafana:10.3.1 Signed-off-by: hagen1778 <roman@victoriametrics.com>	2024-02-05 10:50:36 +02:00
hagen1778	5aa0f77d8c	dashboards: specify where to see details about dropped labels Signed-off-by: hagen1778 <roman@victoriametrics.com>	2024-01-29 17:23:38 +01:00
hagen1778	12acea2584	dashboards: update cluster dashboard * add panels for detailed visualization of traffic usage between vmstorage, vminsert, vmselect components and their clients. New panels are available in the rows dedicated to specific components. * update "Slow Queries" panel to show percentage of the slow queries to the total number of read queries served by vmselect. The percentage value should make it more clear for users whether there is a service degradation. Signed-off-by: hagen1778 <roman@victoriametrics.com> (cherry picked from commit `463455665b`)	2024-01-08 11:58:56 +01:00
Aliaksandr Valialkin	5c43f2261e	dashboards: remove `path!="/favicon.ico"` filter from `requests rate` graphs The `path!="/favicon.ico"` filter has little sense, since there are many other special paths, which may be filtered out - /metrics, /flags, /health, /ping, /robots.txt, /-/healthy, /-/ready, /reload, etc. See /lib/httpserver/httpserver.go for more details. It will be hard or impossible to maintain filters for all these paths, so it is better to drop this filter in order to simplify queries and improve the consistency of these queries.	2023-11-16 19:29:46 +01:00
hagen1778	7d72474a38	dashboards: use `version` instead of `short_version` in annotations `version` label won't show the difference if various flavors of the same version were deployed. But `short_version` will. For example, on the sandbox env we test VM builds before new version release. Without this change, the version update won't be visible on dashboard. Signed-off-by: hagen1778 <roman@victoriametrics.com> (cherry picked from commit `d389a4fcf3`)	2023-11-16 09:27:42 +01:00
hagen1778	72a40539b0	dashboards: update description for RSS and anonymous memory panels to be consistent for single-node, cluster and vmagent dashboards. Signed-off-by: hagen1778 <roman@victoriametrics.com> (cherry picked from commit `d3ae2b2f62`)	2023-11-14 10:00:11 +01:00
hagen1778	777424082b	deployment/dashboards: respect `job` and `instance` filters for `alerts` annotation in cluster and single-node dashboards Signed-off-by: hagen1778 <roman@victoriametrics.com> (cherry picked from commit `d6ae082598`)	2023-11-14 10:00:11 +01:00
hagen1778	8c3bac8f40	dashboards/cluster: fix description about `max` threshold for `Concurrent selects` panel. Before, it was mistakenly implying that `max` is equal to the double of available CPUs. Addresses https://github.com/VictoriaMetrics/VictoriaMetrics/pull/5214 Signed-off-by: hagen1778 <roman@victoriametrics.com>	2023-10-31 19:03:21 +01:00
hagen1778	b57e8b1bb9	dasbhoards: fix vminsert/vmstorage/vmselect metrics filtering Fix vminsert/vmstorage/vmselect metrics filtering when dashboard is used to display data from many sub-clusters with unique job names. Before, only one specific job could have been accounted for component-specific panels, instead of all available jobs for the component. Signed-off-by: hagen1778 <roman@victoriametrics.com>	2023-10-16 12:13:01 +02:00
hagen1778	05b4fbf0b5	dashboards: correctly calculate `Bytes per point` value Correctly calculate `Bytes per point` value for single-server and cluster VM dashboards. Before, the calculation mistakenly accounted for the number of entries in indexdb in denominator, which could have shown lower values than expected. Signed-off-by: hagen1778 <roman@victoriametrics.com>	2023-08-11 04:53:56 -07:00
Zakhar Bessarab	0a8d39e0d5	dashboards/cluster: fix using storage filter for cache usage panel (#4657 ) Using `job=~$job_storage` forces "Cache usage" panel to display only vmstorage caches, but there is a cache peresent at vmselect(`promql/rollupResult`). Updated selector to match generic `$job` so that all caches will be displayed with an option to display per-job caches. Signed-off-by: Zakhar Bessarab <z.bessarab@victoriametrics.com>	2023-07-18 16:00:44 -07:00
Roman Khavronenko	ecd7ec4832	Dashboard upd (#4438 ) dashboards: update dashboard for single-node version * add anonymous mem usage panel; * add syscall rate panel; * add location to logs panel; * update legend for panels to reflect instance name; * update queries to aggregate per instance. dashboards: update dashboard for cluster version * add syscall rate panel; * add drilldown to logs panel. Signed-off-by: hagen1778 <roman@victoriametrics.com>	2023-07-06 16:49:42 -07:00
Aliaksandr Valialkin	531b35b6c0	docs/Troubleshooting.md: document an additional case, which could result in slow inserts If `-cacheExpireDuration` is lower than the interval between ingested samples for the same time series, then vm_slow_row_inserts_total` metric is increased. See https://github.com/VictoriaMetrics/VictoriaMetrics/issues/3976#issuecomment-1476883183	2023-03-20 14:33:27 -07:00
Roman Khavronenko	2b8edaa609	Dashboards upd (#3942 ) * dashboards/cluser: use `quantile` since `median` isn't supported by PromQL Signed-off-by: hagen1778 <roman@victoriametrics.com> * dashboards/: add `restarts` annotation to show when there were restarts The cluster's annotation query is aggregated `by job`, while vmagent/vmalert are aggregated `by job, instance`. This is because cluster dashboard can contains too many instances and annotation could become too noisy. Signed-off-by: hagen1778 <roman@victoriametrics.com> dashboards/*: support instance filter in Version annotation Signed-off-by: hagen1778 <roman@victoriametrics.com> --------- Signed-off-by: hagen1778 <roman@victoriametrics.com>	2023-03-12 00:14:32 -08:00
Roman Khavronenko	c9ee4e5e3d	dashboards: account for indexdb size in Bytes-per-Point panel (#3884 ) Signed-off-by: hagen1778 <roman@victoriametrics.com>	2023-03-08 00:08:41 -08:00
Roman Khavronenko	381dce79e6	dashboards: use `median` instead of `avg` (#3800 ) `avg` can be affected by just one outlier, which may lead to false conclusions. `median` is supposed to reflect reality better by leveling outliers out. Signed-off-by: hagen1778 <roman@victoriametrics.com>	2023-02-11 12:09:09 -08:00
Aliaksandr Valialkin	3db8d7cb01	dashboards: typo fix `Datapoints scanned per series` -> `Datapoints scanned per query`	2023-02-03 19:12:42 -08:00
Roman Khavronenko	17b5ac7e1c	dasbhoards: fix the tooltip info for 1.86 (#3628 ) See `c63755c316 (diff-bba263a473e7fbc9d0fde075ebef6b3d4e32c322ee1210a3e07182292c7723aaR18)` Signed-off-by: hagen1778 <roman@victoriametrics.com>	2023-01-11 16:59:02 -08:00
Aliaksandr Valialkin	b275983403	lib/writeconcurrencylimiter: improve the logic behind -maxConcurrentInserts limit Previously the -maxConcurrentInserts was limiting the number of established client connections, which write data to VictoriaMetrics. Some of these connections could be idle. Such connections do not consume big amounts of CPU and RAM, so there is a little sense in limiting the number of such connections. So now the -maxConcurrentInserts command-line option limits the number of concurrently executed insert requests, not including idle connections. It is recommended removing -maxConcurrentInserts command-line option, since the default value for this option should work good for most cases.	2023-01-06 22:07:16 -08:00
Thomas Danielsson	ec1f6811a1	dashboards: fix operator datasource variable (#3604 ) Got "Failed to upgrade legacy queries Datasource $ds was not found" in Grafana on operator dashboard. It's datasource variable was incorrectly named `datasource`. Also made the rest of the dashboards have homogeneous datasource-variable names and selections, matching vmagent dashboard.	2023-01-05 16:49:19 -08:00
Roman Khavronenko	9b82eebc3e	dashboards: respect $job var in sub-vars for cluster dash (#3487 ) Previously, $job_select, $job_storage and $job_insert didn't respect the $job filter. This change updates the variable queries to account for set $job variable. Signed-off-by: hagen1778 <roman@victoriametrics.com>	2022-12-16 16:44:10 -08:00
Roman Khavronenko	4917c9ad8a	dashboards: add VersionChange annotation (#3473 ) The new annotation is hidden by default and suppose to show component `short_version` label change on the panels. Signed-off-by: hagen1778 <roman@victoriametrics.com>	2022-12-12 14:41:46 -08:00
Aliaksandr Valialkin	d2e34b8052	{dashboards,alerts}: subtitute `{type="indexdb"}` with `{type=~"indexdb.*"}` inside queries after `8189770c50` Updates https://github.com/VictoriaMetrics/VictoriaMetrics/issues/3337	2022-12-05 16:00:42 -08:00
Roman Khavronenko	73340dcb01	dashboards: add `Disk space usage %` and `Disk space usage % by type` panels (#3436 ) The new panels have been added to the vmstorage and drilldown rows. `Disk space usage %` is supposed to show disk space usage percentage. This panel is now also referred by `DiskRunsOutOfSpace` alerting rule. This panel has Drilldown option to show absolute values. `Disk space usage % by type` shows the relation between datapoints and indexdb size. It supposed to help identify cases when indexdb starts to take too much disk space. This panel has Drilldown option to show absolute values. Signed-off-by: hagen1778 <roman@victoriametrics.com>	2022-12-05 00:19:36 -08:00
Roman Khavronenko	51993e7d3e	dashboards: fix typo in data link (#3426 ) Fixes a missing `&` char in data link for ETA panel on cluster dashboards. Without `&` char it generates wrong link when click on Drilldown menu. Signed-off-by: hagen1778 <roman@victoriametrics.com>	2022-12-02 19:05:42 -08:00
Roman Khavronenko	85d0cbbfc6	dashboards: update VM cluster dash (#3401 ) The change list is the following: * bump Grafana version to 9.2.6; * remove artifacts in data links. Signed-off-by: hagen1778 <roman@victoriametrics.com>	2022-11-28 16:43:58 -08:00
Timur Bakeyev	b6064dd645	Update `datasource` entries consistently contain type `prometheus` and uid `$ds`. (#3393 ) Co-authored-by: Timour I. Bakeev <tbakeev@ripe.net>	2022-11-28 16:43:58 -08:00
Roman Khavronenko	ed39d0d11c	dashboards: cleanup & remove artifacts (#3387 ) * some unexpected DS UIDs were removed; * replace `$instance.` filter with `$instance` since we respect the instance port anyway; remove predefined datasource for `clusterbytenant` in favour of datasource variable `ds`. Signed-off-by: hagen1778 <roman@victoriametrics.com> Signed-off-by: hagen1778 <roman@victoriametrics.com>	2022-11-25 07:28:24 -08:00
Roman Khavronenko	cae148d5c6	dashboards: cluster dashboard update (#3380 ) The purpose of the update is to make the dash more usable for large installations with many instances. Panels which showed metrics per-instance (Mem, CPU) now are showing metrics per-job or min/max/avg aggregations in % instead. This supposed to help immediately to identify resource shortage and remain usable for small and big installations. For cases when detailed info is needed, to the bottom of the dashboard a new row `Drilldown` was added. Panels like Mem or CPU now contain a `data-link` named `Drilldown` (cis shown on line click) which takes user to more detailed panel. The change list is the following: * bump Grafana version to 9.1.0; * replace old "Graph" panel with "TimeSeries" panel; * improve Uptime panel to show number of instances per job; * show % usage of Mem and CPU instead of absolute values; * `Caches` row was removed. All needed info for caches is now part of `Troubleshooting`; * add `Drilldown` section for detailed resource usage; * add Annotations for Alert triggers. Not all alerts are supposed to be displayed on the dashboard, but only those with label `show_at: dashboard`. See `alerts-cluster.yml` change. Signed-off-by: hagen1778 <roman@victoriametrics.com> Signed-off-by: hagen1778 <roman@victoriametrics.com>	2022-11-24 13:20:10 -08:00
Roman Khavronenko	0efc20d7b8	dashboards: replace `Index size` panel with `Active series` (#3157 ) Panel `Index size` showed itself impractical for users. So replacing it with `Active series` panel. https://github.com/VictoriaMetrics/VictoriaMetrics/issues/776#issuecomment-1255823734 Signed-off-by: hagen1778 <roman@victoriametrics.com>	2022-09-26 08:48:25 +03:00
Roman Khavronenko	5dfe63e102	Dashboards (#3120 ) * dashboards/cluster: few updates * apply consistent formatting across panels; * make resource usage panels per component more detailed; * add extra panels to vmselect for displaying `vm_rows_read_per_query`, `vm_rows_scanned_per_query`, `vm_rows_read_per_series` and `vm_series_read_per_query` metrics. Signed-off-by: hagen1778 <roman@victoriametrics.com> * dashboards/single: few updates * apply consistent formatting across panels; * add extra panels to Performance for displaying `vm_rows_read_per_query`, `vm_rows_scanned_per_query`, `vm_rows_read_per_series` and `vm_series_read_per_query` metrics. Signed-off-by: hagen1778 <roman@victoriametrics.com> * dashboards/vmagent: few updates * apply consistent formatting across panels; * add panels for showing number of samples ingested or scraped; * adapt resource usage panels for multiple selected jobs/instances; * add adhoc variable; * display vmagent's version in Stats. Signed-off-by: hagen1778 <roman@victoriametrics.com> * dashboards/vmalert: few updates * apply consistent formatting across panels; * adapt resource usage panels for multiple selected jobs/instances; * show vmalert version in Stats section. Signed-off-by: hagen1778 <roman@victoriametrics.com>	2022-09-19 15:04:37 +03:00
Max Golionko	e07f23a1b9	moved cluster dashboard to master (#3074 ) dashboards: move cluster dashboard to master branch This change should simplify dashboards management.	2022-09-08 11:47:25 +03:00

48 Commits