prometheus

mirror of https://github.com/prometheus/prometheus.git synced 2025-02-21 03:16:00 -08:00

History

beorn7 fc6737b7fb storage: improve index lookups tl;dr: This is not a fundamental solution to the indexing problem (like tindex is) but it at least avoids utilizing the intersection problem to the greatest possible amount. In more detail: Imagine the following query: nicely:aggregating:rule{job="foo",env="prod"} While it uses a nicely aggregating recording rule (which might have a very low cardinality), Prometheus still intersects the low number of fingerprints for `{__name__="nicely:aggregating:rule"}` with the many thousands of fingerprints matching `{job="foo"}` and with the millions of fingerprints matching `{env="prod"}`. This totally innocuous query is dead slow if the Prometheus server has a lot of time series with the `{env="prod"}` label. Ironically, if you make the query more complicated, it becomes blazingly fast: nicely:aggregating:rule{job=~"foo",env=~"prod"} Why so? Because Prometheus only intersects with non-Equal matchers if there are no Equal matchers. That's good in this case because it retrieves the few fingerprints for `{__name__="nicely:aggregating:rule"}` and then starts right ahead to retrieve the metric for those FPs and checking individually if they match the other matchers. This change is generalizing the idea of when to stop intersecting FPs and go into "retrieve metrics and check them individually against remaining matchers" mode: - First, sort all matchers by "expected cardinality". Matchers matching the empty string are always worst (and never used for intersections). Equal matchers are in general consider best, but by using some crude heuristics, we declare some better than others (instance labels or anything that looks like a recording rule). - Then go through the matchers until we hit a threshold of remaining FPs in the intersection. This threshold is higher if we are already in the non-Equal matcher area as intersection is even more expensive here. - Once the threshold has been reached (or we have run out of matchers that do not match the empty string), start with "retrieve metrics and check them individually against remaining matchers". A beefy server at SoundCloud was spending 67% of its CPU time in index lookups (fingerprintsForLabelPairs), serving mostly a dashboard that is exclusively built with recording rules. With this change, it spends only 35% in fingerprintsForLabelPairs. The CPU usage dropped from 26 cores to 18 cores. The median latency for query_range dropped from 14s to 50ms(!). As expected, higher percentile latency didn't improve that much because the new approach is _occasionally_ running into the worst case while the old one was _systematically_ doing so. The 99th percentile latency is now about as high as the median before (14s) while it was almost twice as high before (26s).		2016-07-20 17:35:53 +02:00
..
codable	Replace metric.LabelPair with model.LabelPair	2015-08-22 13:32:13 +02:00
fixtures/b0	Add benchmark for loading chunks and chunk descs.	2015-03-19 19:28:21 +01:00
index	Handle errors caused by data corruption more gracefully	2016-03-02 23:02:34 +01:00
storagetool	Make version informations consistent between prometheus components	2016-05-05 22:33:18 +02:00
chunk.go	storage: Make MemorySeriesStorage a public type	2016-06-29 08:14:23 +02:00
crashrecovery.go	Crash recovery: Fix an edge case.	2016-07-07 16:17:38 +02:00
delta.go	Implement Gorilla-inspired chunk encoding	2016-03-17 14:47:08 +01:00
delta_helpers.go	Switch from client_golang/model to common/model	2015-08-21 13:33:38 +02:00
doubledelta.go	Implement Gorilla-inspired chunk encoding	2016-03-17 14:47:08 +01:00
heads.go	Merge branch 'master' into beorn7/storage4	2016-03-08 00:14:00 +01:00
instrumentation.go	storage: Make MemorySeriesStorage a public type	2016-06-29 08:14:23 +02:00
interface.go	storage: improve index lookups	2016-07-20 17:35:53 +02:00
locker.go	Update doc comments	2016-06-03 12:34:01 +02:00
locker_test.go	Add missing license headers	2016-04-13 16:08:22 +02:00
mapper.go	Checkpoint fingerprint mappings only upon shutdown	2016-04-15 01:03:28 +02:00
mapper_test.go	Checkpoint fingerprint mappings only upon shutdown	2016-04-15 01:03:28 +02:00
persistence.go	Consistently use the `Seconds()` method for conversion of durations	2016-07-07 15:24:35 +02:00
persistence_test.go	Implement Gorilla-inspired chunk encoding	2016-03-17 14:47:08 +01:00
preload.go	storage: Make MemorySeriesStorage a public type	2016-06-29 08:14:23 +02:00
series.go	storage: Make MemorySeriesStorage a public type	2016-06-29 08:14:23 +02:00
series_test.go	Never drop a still open head chunk.	2016-04-15 19:18:40 +02:00
storage.go	storage: improve index lookups	2016-07-20 17:35:53 +02:00
storage_test.go	storage: improve index lookups	2016-07-20 17:35:53 +02:00
test_helpers.go	storage: Make MemorySeriesStorage a public type	2016-06-29 08:14:23 +02:00
varbit.go	Work around compiler bug	2016-03-29 17:05:28 +02:00
varbit_helpers.go	Rename Gorilla into varbit	2016-03-23 16:30:41 +01:00
varbit_test.go	Rename Gorilla into varbit	2016-03-23 16:30:41 +01:00