prometheus

mirror of https://github.com/prometheus/prometheus.git synced 2025-03-05 20:59:13 -08:00

Author	SHA1	Message	Date
Carrie Edwards	45a32a29ef	Update tsdb tests to use test utils. Co-authored-by: Fiona Liao <fiona.liao@grafana.com> Signed-off-by: Carrie Edwards <edwrdscarrie@gmail.com>	2024-07-03 09:28:38 -07:00
Carrie Edwards	60917f628b	Add test utilities for testing different sample types. Co-authored-by: Fiona Liao <fiona.liao@grafana.com> Signed-off-by: Carrie Edwards <edwrdscarrie@gmail.com>	2024-07-03 09:28:38 -07:00
Bryan Boreham	134e8dc7af	TSDB: Simplify OOO Select by copying the head chunk (#14396 ) Instead of carrying around extra fields in `Meta` structs which let us approximate what was in the chunk at the time, take a copy of the chunk. This simplifies lots of code, and lets us correct a couple of tests which were embedding the wrong answer. We can also remove boundedIterator, which was only used to constrain the OOO head chunk. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-07-03 15:08:07 +01:00
Oleg Zaytsev	726ed124e4	Replace `ListPostings.Seek`'s binary search call by the generic `slices.BinarySearch` (#14393 ) Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2024-07-02 14:51:05 +01:00
Arve Knudsen	c3793f2b32	Merge remote-tracking branch 'prometheus/main' into arve/wlog-histograms Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2024-07-01 10:51:22 +02:00
Björn Rabenstein	2e58d46522	Merge pull request #13662 from prometheus/nhcb Native histograms custom buckets storage	2024-06-27 21:44:20 +02:00
Bryan Boreham	348f7f8d0c	Merge pull request #14341 from charleskorn/charleskorn/cleanup-pending-read Fix issue where pending OOO read can be left dangling if creating querier fails	2024-06-25 09:23:54 +01:00
Ben Ye	246b7c6a5c	TSDB: Change block populator to accept postings index function (#14213 ) Signed-off-by: Ben Ye <benye@amazon.com>	2024-06-25 09:21:48 +01:00
Ben Ye	5585a3c7e5	tsdb: expose hook to customize block querier (#14114 ) * expose hook for block querier Signed-off-by: Ben Ye <benye@amazon.com> * update comment Signed-off-by: Ben Ye <benye@amazon.com> * use defined type Signed-off-by: Ben Ye <benye@amazon.com> --------- Signed-off-by: Ben Ye <benye@amazon.com>	2024-06-25 09:47:06 +02:00
Charles Korn	2c5e88748e	Fix issue where pending OOO read can be left dangling if creating querier fails Signed-off-by: Charles Korn <charles.korn@grafana.com>	2024-06-25 14:22:44 +10:00
🌲 Harry 🌊 John 🏔	d5f6887294	Pass limit param as hint to storage.Querier Signed-off-by: 🌲 Harry 🌊 John 🏔 <johrry@amazon.com>	2024-06-20 09:47:38 -07:00
Jeanette Tan	dda5f48c9e	Merge branch 'main' into nhcb-review-2	2024-06-20 22:50:00 +08:00
Oleg Zaytsev	fd1a89b7c8	Pass affected labels to `MemPostings.Delete()` (#14307 ) * Pass affected labels to MemPostings.Delete As suggested by @bboreham, we can track the labels of the deleted series and avoid iterating through all the label/value combinations. This looks much faster on the MemPostings.Delete call. We don't have a benchmark on stripeSeries.gc() where we'll pay the price of iterating the labels of each one of the deleted series. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2024-06-18 10:28:56 +00:00
Ben Ye	0e6fca8e76	add unit test Signed-off-by: Ben Ye <benye@amazon.com>	2024-06-16 12:09:42 -07:00
Ben Ye	e7db2e30a4	fix check context cancellation not incrementing count Signed-off-by: Ben Ye <benye@amazon.com>	2024-06-15 11:43:26 -07:00
Ben Ye	5a218708f1	tsdb: Extend compactor interface to allow compactions to create multiple output blocks (#14143 ) * add hook to allow head compaction to create multiple output blocks Signed-off-by: Ben Ye <benye@amazon.com> * change Compact interface; remove BlockPopulator changes Signed-off-by: Ben Ye <benye@amazon.com> * rebase main Signed-off-by: Ben Ye <benye@amazon.com> * fix lint Signed-off-by: Ben Ye <benye@amazon.com> * fix unit test Signed-off-by: Ben Ye <benye@amazon.com> * address feedbacks; add unit test Signed-off-by: Ben Ye <benye@amazon.com> * Apply suggestions from code review Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Update tsdb/compact_test.go Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> --------- Signed-off-by: Ben Ye <benye@amazon.com> Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com>	2024-06-12 17:31:25 -04:00
Sebastian Rabenhorst	05380aa0ac	agent db: make rejecting ooo samples configurable (#14094 ) feat: Make OOO ingestion time window configurable for Prometheus Agent. Signed-off-by: Sebastian Rabenhorst <sebastian.rabenhorst@shopify.com>	2024-06-12 11:07:42 -03:00
Oleg Zaytsev	64a9abb8be	Change LabelValuesFor() to accept index.Postings (#14280 ) The only call we have to LabelValuesFor() has an index.Postings, and we expand it to pass to this method, which will iterate over the values. That's a waste of resources: we can iterate on the index.Postings directly. If there's any downstream implementation that has a slice of series, they can always do an index.ListPostings from them: doing that is cheaper than expanding an abstract index.Postings. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2024-06-11 15:36:46 +02:00
Bryan Boreham	c5d923aa7c	Merge pull request #14279 from colega/fix-label-names-for-not-found headIndexReader.LabelNamesFor: skip not found series	2024-06-11 01:06:19 +03:00
Oleg Zaytsev	10a3c7220b	`MemPostings.PostingsForLabelMatching()`: don't hold the mutex while matching (#14286 ) * MemPostings.PostingsForLabelMatching: let mutex go This changes the `MemPostings.PostingsForLabelMatching` implementation to stop holding the read mutex while matching the label values. We've seen that this method can be slow when the matcher is expensive, that's why we even added a context expiration check. However, there are critical process that might be waiting on this mutex: writes (adding new series) and compaction (deleting the garbage-collected ones), so we should avoid holding it for a long period of time. Given that we've copied the values to a slice anyway, there's no need to hold the lock while matching. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2024-06-10 14:24:17 +02:00
Oleg Zaytsev	2dc177d8af	`MemPostings.Delete()`: reduce locking/unlocking (#13286 ) * MemPostings: reduce locking/unlocking MemPostings.Delete is called from Head.gc(), i.e. it gets the IDs of the series that have churned. I'd assume that many label values aren't affected by that churn at all, so it doesn't make sense to touch the lock while checking them. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2024-06-10 14:23:22 +02:00
Oleg Zaytsev	d0d361da53	headIndexReader.LabelNamesFor: skip not found series It's quite common during the compaction cycle to hold series IDs for series that aren't in the TSDB head anymore. We shouldn't fail if that happens, as the caller has no way to figure out which one of the IDs doesn't exist. Fixes https://github.com/prometheus/prometheus/issues/14278 Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2024-06-07 16:09:53 +02:00
Jeanette Tan	14f8dded39	Merge branch 'main' into nhcb Signed-off-by: Jeanette Tan <jeanette.tan@grafana.com>	2024-06-07 19:17:14 +08:00
Jeanette Tan	9adc1699c3	fix according to code review Signed-off-by: Jeanette Tan <jeanette.tan@grafana.com>	2024-06-07 18:50:59 +08:00
Ben Ye	8a08f452b6	tsdb: Allow passing a custom compactor to override the default one (#14113 ) * expose hook in tsdb to allow customizing compactor Signed-off-by: Ben Ye <benye@amazon.com> * address comment Signed-off-by: Ben Ye <benye@amazon.com> --------- Signed-off-by: Ben Ye <benye@amazon.com>	2024-06-04 19:11:36 -04:00
Bryan Boreham	42b546a43d	tsdb: add details to duplicate sample error (#13277 ) Now the error will include the timestamp and the existing and new values. When you are trying to track down the source of this error, it can be useful to see that the values are close, or alternating, or something else. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-06-04 08:54:09 +01:00
Arve Knudsen	b8b9015e38	tsdb/index: Fix TestReader_PostingsForLabelMatchingHonorsContextCancel Fix number of series in TestReader_PostingsForLabelMatchingHonorsContextCancel (off by one). Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2024-06-03 17:29:06 +02:00
Bryan Boreham	3ee52abb53	[ENHANCEMENT] TSDB: Save map lookup on validation Goes faster. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-05-30 09:17:11 +01:00
Bryan Boreham	7d98487447	[ENHANCEMENT] TSDB: let Resize re-use buffer This saves having to zero the buffer every time. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-05-30 09:17:11 +01:00
Bryan Boreham	c0bb156eca	[ENHANCEMENT] TSDB: Eliminate pointer when storing exemplars Saves memory and effort. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-05-30 09:17:11 +01:00
Bryan Boreham	3eb5581877	[ENHANCEMENT] TSDB: Reduce map lookups on exemplar index In many cases we already have a pointer to the entry. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-05-30 09:17:11 +01:00
Bryan Boreham	f0c50b5a66	[Test] TSDB: BenchmarkResizeExemplar multiple per series One exemplar per series is not a typical workload. Make it the same as `BenchmarkAddExemplar`. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-05-30 09:17:11 +01:00
Bryan Boreham	929fbf860e	[Test] TSDB: let BenchmarkAddExemplar reuse slots Test with different amounts of capacity and exemplars, so that sometimes new exemplars are evicting older exemplars. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-05-30 09:16:30 +01:00
Ben Ye	6683895620	optimize regex matching for empty label values in posting match (#14075 ) Also update tests. Signed-off-by: Ben Ye <benye@amazon.com>	2024-05-29 16:03:33 +01:00
Arve Knudsen	694f717dc4	Watcher.readSegment: Only consider unknown rec types failures Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2024-05-28 15:23:50 +02:00
Arve Knudsen	a4942ffa8c	Merge remote-tracking branch 'prometheus/main' into arve/wlog-histograms	2024-05-28 12:06:29 +02:00
Arve Knudsen	b2396c0c8f	Upgrade to golangci-lint v1.59.0 Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2024-05-27 22:38:48 +02:00
Arve Knudsen	90145fa884	Merge remote-tracking branch 'prometheus/main' into arve/wlog-histograms	2024-05-27 15:19:33 +02:00
Alan Protasio	8894d65cd6	Fix head stats and hooks when replaying a corrupted snapshot (#14079 ) * Fixing head stats and hooks when replaying a corrupted snapshot Signed-off-by: alanprot <alanprot@gmail.com> * Fixing create/removed series metrics Signed-off-by: alanprot <alanprot@gmail.com> * Refactoring to have common code between gc and flush method Signed-off-by: alanprot <alanprot@gmail.com> * Update tsdb/head.go Co-authored-by: Ayoub Mrini <ayoubmrini424@gmail.com> Signed-off-by: Alan Protasio <alanprot@gmail.com> * refactor Signed-off-by: alanprot <alanprot@gmail.com> * Update tsdb/head_test.go Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com> Signed-off-by: Alan Protasio <alanprot@gmail.com> * Update tsdb/head_test.go Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com> Signed-off-by: Alan Protasio <alanprot@gmail.com> --------- Signed-off-by: alanprot <alanprot@gmail.com> Signed-off-by: Alan Protasio <alanprot@gmail.com> Co-authored-by: Ayoub Mrini <ayoubmrini424@gmail.com> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com>	2024-05-24 22:43:21 -04:00
Björn Rabenstein	3119b8a055	Merge pull request #13218 from machine424/ro-promtool Make DBReadOnly more RO	2024-05-21 13:27:40 +02:00
Oleg Zaytsev	fe9cb5a803	Check context every 128 labels instead of 100 (#14118 ) Follow up on https://github.com/prometheus/prometheus/pull/14096 As promised, I bring a benchmark, which shows a very small improvement if context is checked every 128 iterations of label instead of every 100. It's much easier for a computer to check modulo 128 than modulo 100. This is a very small 0-2% improvement but I'd say this is one of the hottest paths of the app so this is still relevant. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2024-05-21 11:30:43 +02:00
Arve Knudsen	5ca56eeb6b	tsdb/index: Refactor Reader tests (#14071 ) tsdb/index: Refactor Reader tests Co-authored-by: Björn Rabenstein <github@rabenste.in> Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com> --------- Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com> Co-authored-by: Björn Rabenstein <github@rabenste.in>	2024-05-16 11:51:46 +02:00
Oleksandr Redko	f10c3454e9	Enable perfsprint linter and fix up code Signed-off-by: Oleksandr Redko <oleksandr.red+github@gmail.com>	2024-05-15 17:51:05 +03:00
György Krajcsovits	b215a41be4	tsdb/index/postings: fix missing lock unlock Followup to #14096 Unfortunately the previous PR introduced this bug by not releasing the lock before returning. Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com>	2024-05-15 14:02:39 +02:00
George Krajcsovits	fdaafdb041	tsdb: check for context cancel before regex matching postings (#14096 ) * tsdb: check for context cancel before regex matching postings Regex matching can be heavy if the regex takes a lot of cycles to evaluate and we can get stuck evaluating postings for a long time without this fix. The constant checkContextEveryNIterations=100 may be changed later. Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com>	2024-05-15 06:26:19 +02:00
Jeanette Tan	f028496133	Merge branch 'main' into nhcb Signed-off-by: Jeanette Tan <jeanette.tan@grafana.com>	2024-05-14 16:20:15 +08:00
Arve Knudsen	f094dd50e0	Merge remote-tracking branch 'prometheus/main' into arve/wlog-histograms	2024-05-09 11:58:37 +02:00
Arve Knudsen	5c4310aa37	[ENHANCEMENT] TSDB: Optimize querying with regexp matchers Add method `PostingsForLabelMatching` to `tsdb.IndexReader`, to obtain postings for labels with a certain name and values accepted by a provided callback, and use it from `tsdb.PostingsForMatchers`. The intention is to optimize regexp matcher paths, especially not having to load all label values before matching on them. Plus tests, and refactor some `tsdb/index.Reader` methods. Benchmarking shows memory reduction up to ~100%, and speedup of up to ~50%. Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com> Co-authored-by: Bartlomiej Plotka <bwplotka@gmail.com>	2024-05-09 10:55:30 +01:00
Arve Knudsen	90f08de96a	Merge remote-tracking branch 'prometheus/main' into arve/wlog-histograms	2024-05-08 18:14:30 +02:00
Arve Knudsen	d699dc3c77	Fix language in docs and comments (#14041 ) Fix language in docs and comments --------- Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com> Co-authored-by: Björn Rabenstein <github@rabenste.in>	2024-05-08 17:57:09 +02:00
Arve Knudsen	f1e4c33943	Merge remote-tracking branch 'prometheus/main' into arve/wlog-histograms	2024-05-08 15:06:41 +02:00
Arve Knudsen	108a6bc9f6	tsdb/chunkenc.Pool: Refactor Get and Put Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2024-05-08 13:37:25 +02:00
Jeanette Tan	796b1bbfde	Merge branch 'main' into nhcb Signed-off-by: Jeanette Tan <jeanette.tan@grafana.com>	2024-05-08 19:11:39 +08:00
Arve Knudsen	fb6a45f06b	tsdb/wlog: Only treat unknown record types as failure Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2024-05-06 11:58:02 +02:00
Alan Protasio	d15869af32	Avoid creating new slices for labels values on postings for matchers (#13958 ) * Avoid creating new slices for labels values on postings for matchers Signed-off-by: alanprot <alanprot@gmail.com> * refactor Signed-off-by: alanprot <alanprot@gmail.com> --------- Signed-off-by: alanprot <alanprot@gmail.com>	2024-04-24 16:41:33 +02:00
György Krajcsovits	bcafa5f1f9	Merge remote-tracking branch 'upstream/main' into update-nhcb	2024-04-24 11:06:59 +02:00
Arthur Silva Sens	b5b5e1e5ae	Merge pull request #13919 from GiedriusS/dont_forget_to_unregister tsdb/wlog: unregister metrics on WL close	2024-04-18 16:44:03 -03:00
Giedrius Statkevičius	bdf490726a	tsdb/wlog: add test for metrics unregistering Signed-off-by: Giedrius Statkevičius <giedrius.statkevicius@vinted.com>	2024-04-18 11:11:37 +03:00
machine424	c5a1cc9148	chore(tsdb): add a sandboxDir to DBReadOnly, the directory can be used for transient file writes. use it in loadDataAsQueryable to make sure the RO Head doesn't truncate or cut new chunks in data/chunks_head/. add a -sandbox-dir-root flag to "promtool tsdb dump/dump-openmetrics" to control the root of that sandbox dirrectory. Signed-off-by: machine424 <ayoubmrini424@gmail.com>	2024-04-15 17:00:25 +02:00
Giedrius Statkevičius	3b8fe00767	tsdb/wlog: unregister metrics on WL close Thanos can create and destroy TSDBs dynamically, and once a TSDB disappears its files are deleted. Calculating the size of the WAL then fails with errors like: ``` msg: "Failed to calculate size of "wal" dir", "err": "lstat /tsdbdir/wal: no such file or directory", "caller": "wlog.go:271" ``` Signed-off-by: Giedrius Statkevičius <giedrius.statkevicius@vinted.com>	2024-04-11 11:30:05 +03:00
Matthieu MOREL	6f595c6762	golangci-lint: enable whitespace linter (#13905 ) Signed-off-by: Matthieu MOREL <matthieu.morel35@gmail.com>	2024-04-11 09:27:54 +01:00
Jonathan Halterman	633224886a	Write out of order hint when initially creating meta file (#13894 ) Signed-off-by: Jonathan Halterman <jonathan@grafana.com> Signed-off-by: Jonathan Halterman <jhalterman@gmail.com> Co-authored-by: Jesus Vazquez <jesusvazquez@users.noreply.github.com>	2024-04-08 17:34:14 +02:00
Łukasz Mierzwa	277f04f0c4	Stop compactions if there's a block to write (#13754 ) * Stop compactions if there's a block to write db.Compact() checks if there's a block to write with HEAD chunks before calling db.compactBlocks(). This is to ensure that if we need to write a block then it happens ASAP, otherwise memory usage might keep growing. But what can also happen is that we don't need to write any block, we start db.compactBlocks(), compaction takes hours, and in the meantime HEAD needs to write out chunks to a block. This can be especially problematic if, for example, you run Thanos sidecar that's uploading block, which requires that compactions are disabled. Then you disable Thanos sidecar and re-enable compactions. When db.compactBlocks() is finally called it might have a huge number of blocks to compact, which might take a very long time, during which HEAD cannot write out chunks to a new block. In such case memory usage will keep growing until either: - compactions are finally finished and HEAD can write a block - we run out of memory and Prometheus gets OOM-killed This change adds a check for pending HEAD block writes inside db.compactBlocks(), so that we bail out early if there are still compactions to run, but we also need to write a new block. Also add a test for compactBlocks. --------- Signed-off-by: Łukasz Mierzwa <l.mierzwa@gmail.com> Signed-off-by: Lukasz Mierzwa <lukasz@cloudflare.com>	2024-04-07 18:28:28 +01:00
Bryan Boreham	fc567a1df8	Merge pull request #13889 from komisan19/refactor/utilize_standard_functions_max/min refactor: utilize standard functions max/min in promtool and tsdb	2024-04-06 10:23:18 +01:00
Julien	8eb9228af8	Merge pull request #13864 from yeya24/expose-compactor-metrics Expose compactor metrics	2024-04-05 11:24:41 +02:00
Jonathan Halterman	113938aeb8	Log out of order when writing a block (#13888 ) Signed-off-by: Jonathan Halterman <jonathan@grafana.com>	2024-04-04 14:26:13 +02:00
komisan19	0249e080b4	refactor: utilize standard functions max/min Signed-off-by: komisan19 <18901496+komisan19@users.noreply.github.com>	2024-04-04 03:15:38 +09:00
Nicolas Takashi	8125634086	[refactor] moving mergedOOOChunks Iterator (#13881 ) Signed-off-by: Nicolas Takashi <nicolas.tcs@hotmail.com>	2024-04-03 10:14:34 +02:00
carehabit	a672662073	all: fix some typos (#13863 ) Signed-off-by: carehabit <shenyuting@outlook.com>	2024-04-01 18:06:05 +02:00
Ben Ye	ded35ef20d	expose compactor metrics Signed-off-by: Ben Ye <benye@amazon.com>	2024-03-31 15:10:29 -07:00
Nicolas Takashi	0b762db154	[refactor] moving mergedOOOChunks to ooo_head_read Signed-off-by: Nicolas Takashi <nicolas.tcs@hotmail.com>	2024-03-29 23:33:15 +00:00
György Krajcsovits	2a4aa085d2	Merge branch 'main' into nhcb	2024-03-27 18:42:10 +01:00
George Krajcsovits	4eab18abd6	[nhcb branch] Use single bit to differentiate between optimized bounds and floats (#13828 ) * Use single bit to differentiate between optimized bounds and floats Use one bit to decide what kind of data to read/write. This reduces storage need of floats from 72 bits to 65 bits and makes the integers store in 5 to 32 bits instead of 16. Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com> Signed-off-by: Jeanette Tan <jeanette.tan@grafana.com> Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com> Signed-off-by: Jeanette Tan <jeanette.tan@grafana.com> Signed-off-by: George Krajcsovits <krajorama@users.noreply.github.com> Co-authored-by: Jeanette Tan <jeanette.tan@grafana.com>	2024-03-27 18:40:59 +01:00
Arve Knudsen	35aab01de0	tsdb/wlog.Checkpoint: Handle also float histograms Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2024-03-26 15:27:32 +01:00
Bartlomiej Plotka	de05d5e11b	Merge pull request #13776 from aknuds1/arve/checkpoint-histogram-samples tsdb/wlog.Checkpoint: Fix counting of histogram samples	2024-03-26 12:33:48 +01:00
Bartlomiej Plotka	25578f2b22	[test] Merge pull request #13790 from aknuds1/arve/retention-commit tsdb.BeyondTimeRetention: Fix comment and test at retention duration	2024-03-26 12:26:32 +01:00
Nick Pillitteri	481f14e1c0	TSDB: Don't rely on integer overflow in head compaction check (#13755 ) * TSDB: Don't compact the head block when empty Don't compact the Head block if there have not yet been any samples appended. Previously, the logic for determining if the head should be compacted relied on the default values for min and max time and integer overflow when they were checked in `Head.compactable()`. The check in `Head.compactable()` effectively did `math.MinInt64 - math.MaxInt64` which overflowed and wrapped to `1`. Since `1` is less than `1.5` times the chunk range, compaction did not happen. This was the correct behavior but relying on overflow wrapping is surprising. This change add a method for checking if the min and max time for the head is unset and uses it to short-circuit compaction in that case. It also replaces several explicit checks for the default value to determine if the head has not yet had any samples added. Signed-off-by: Nick Pillitteri <nick.pillitteri@grafana.com>	2024-03-26 12:17:38 +01:00
Ben Ye	ceca6c4716	[ENHANCEMENT] TSDB: Log more statistics during startup (#13838 ) * log chunk snapshot and mmap chunks replay duration together with total replay duration Signed-off-by: Ben Ye <benye@amazon.com>	2024-03-26 11:16:27 +00:00
Bryan Boreham	78c0fd2f4d	Merge pull request #13799 from machine424/wbl chore(tsdb): set the wbl to nil as well in DBReadOnly.loadDataAsQuery…	2024-03-25 15:03:56 +01:00
Bryan Boreham	080d440bf8	Merge remote-tracking branch 'origin/main' into pr/13461	2024-03-25 12:14:26 +00:00
György Krajcsovits	a3d1a46eda	Merge branch 'main' into nhcb	2024-03-22 14:51:48 +01:00
zenador	4acbb7dea6	Add custom buckets to native histogram chunks encoding (#13706 ) * add custom bounds to chunks encoding * change custom buckets schema number * rename custom bounds to custom values Signed-off-by: Jeanette Tan <jeanette.tan@grafana.com>	2024-03-22 14:36:39 +01:00
machine424	2a2e2ed28b	chore(tsdb): set the wbl to nil as well in DBReadOnly.loadDataAsQueryable Signed-off-by: machine424 <ayoubmrini424@gmail.com>	2024-03-20 17:13:39 +01:00
Arve Knudsen	07332f7427	TestTimeRetention: Split into two sub-tests Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2024-03-20 15:06:36 +01:00
Arve Knudsen	af694dc295	Merge TestDB_BeyondTimeRetention into TestTimeRetention Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2024-03-20 09:07:16 +01:00
Arve Knudsen	9c7a734063	tsdb.BeyondTimeRetention: Fix comment and test at retention duration Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2024-03-19 09:10:21 +01:00
Bryan Boreham	a0e93e403e	Merge pull request #13764 from bboreham/remove-deprecated-wal [Cleanup] TSDB: Remove old deprecated WAL implementation Deprecated since 2018.	2024-03-17 09:34:57 +00:00
Darshan Chaudhary	b7047f7fcb	Fix retention boundary so 2h retention deletes blocks right at the 2h boundary (#9633 ) Signed-off-by: darshanime <deathbullet@gmail.com>	2024-03-15 19:35:16 +01:00
Arve Knudsen	cef1025ea8	tsdb/wlog.Checkpoint: Fix counting of histogram samples Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2024-03-15 10:23:59 +01:00
Bryan Boreham	d45b5deb75	TSDB: move function only used in tests Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-03-15 08:54:47 +00:00
Bryan Boreham	3274cac0d3	TSDB: remove unused function Was only used in old WAL implementation. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-03-15 08:51:57 +00:00
Arve Knudsen	1de49d5b69	Remove unused function tsdb/chunks.PopulatedChunk (#13763 ) Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2024-03-14 11:15:17 +01:00
Bryan Boreham	87edf1f960	[Cleanup] TSDB: Remove old deprecated WAL implementation Deprecated since 2018. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-03-13 15:57:23 +00:00
Bryan Boreham	d08f054950	[ENHANCEMENT] TSDB: Check CRC without allocating (#13742 ) Use the existing utility function which does this. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-03-12 12:24:27 +01:00
carrychair	856f6e49c8	fix function and struct name Signed-off-by: carrychair <linghuchong404@gmail.com>	2024-03-09 17:53:17 +08:00
Björn Rabenstein	9acae57937	Merge pull request #13681 from krajorama/native-latency-histograms Add native histograms to latency/duration metrics	2024-03-07 20:46:43 +01:00
Bryan Boreham	bbe39af99f	tsdb: zero out Labels and memSeries pointers in pool (#13712 ) * tsdb: zero out Labels and memSeries pointers in pool So that the garbage-collector doesn't see this memory as still in use. Signed-off-by: Bryan Boreham <bjboreham@gmail.com> --------- Signed-off-by: Bryan Boreham <bjboreham@gmail.com> Signed-off-by: Björn Rabenstein <github@rabenste.in> Co-authored-by: Björn Rabenstein <github@rabenste.in>	2024-03-06 13:36:34 +01:00
György Krajcsovits	4d4d822c36	Add native histograms to latency/duration metrics Dogfood native histograms. Allow dependent projects to migrate to native histograms. I took the defaults from client_golang. Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com>	2024-03-01 14:44:38 +01:00
machine424	f477e0539a	Move from golang.org/x/exp/slices into slices now that we only support Go >= 1.21 Prevent adding back golang.org/x/exp/slices. Signed-off-by: machine424 <ayoubmrini424@gmail.com>	2024-02-28 14:54:53 +01:00
Bryan Boreham	925134e6de	tsdb tests: make work with labels SymbolTable Need to initialize decoders with SymbolTable. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-02-26 11:45:25 +00:00
Bryan Boreham	93b72ec5dd	tsdb: create SymbolTables for labels as required Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-02-26 11:45:25 +00:00
Bryan Boreham	4d6bb2e0e4	Merge pull request #13628 from bboreham/cleanup-13583 tsdb/wlog: small cleanup of WAL watcher after #13583	2024-02-23 12:39:17 +00:00
George Krajcsovits	c6c8f63516	Merge pull request #13607 from fionaliao/ooo-samples-appended-type Add sample type label to outOfOrderSamplesAppended metric	2024-02-23 09:41:45 +01:00
Bryan Boreham	6ed56c9f04	WAL watcher: improve comments Clarify in the first comment that it is `watch()` that waits, and reduce verbiage. The second comment was slightly contradictory to the first and otherwise didn't seem to add much, since `currentSegment` was incremented just a few lines later. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-02-22 09:32:46 +00:00
Bryan Boreham	a975a83079	tsdb: clean up Watcher debug messages Print lastSegment after it gets initialized. Move variable declaration to first use. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-02-22 09:19:18 +00:00
Bryan Boreham	78f46bccca	tsdb/wlog tests: remove unnecessary sleep check Sleep() is documented to return immediately on negative or zero input. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-02-22 09:14:52 +00:00
Callum Styan	0c71230784	fix bug that would cause us to endlessly fall behind (#13583 ) * fix bug that would cause us to only read from the WAL on the 15s fallback timer if remote write had fallen behind and is no longer reading from the WAL segment that is currently being written to Signed-off-by: Callum Styan <callumstyan@gmail.com> * remove unintended logging, fix lint, plus allow test to take slightly longer because cloud CI Signed-off-by: Callum Styan <callumstyan@gmail.com> * address review feedback Signed-off-by: Callum Styan <callumstyan@gmail.com> * fix watcher sleeps in test, flu brain is smooth Signed-off-by: Callum Styan <callumstyan@gmail.com> * increase timeout, unfortunately cloud CI can require a longer timeout Signed-off-by: Callum Styan <callumstyan@gmail.com> --------- Signed-off-by: Callum Styan <callumstyan@gmail.com>	2024-02-21 17:09:07 -08:00
Fiona Liao	841a133514	Move histogramsAppended to be more consistent Signed-off-by: Fiona Liao <fiona.liao@grafana.com>	2024-02-21 11:15:04 +00:00
Fiona Liao	52389647b2	Add type label to outOfOrderSamplesAppended metric Signed-off-by: Fiona Liao <fiona.liao@grafana.com>	2024-02-19 15:24:39 +00:00
Bryan Boreham	c0e36e6bb3	Standardise exemplar label as "trace_id" This is consistent with the OpenTelemetry standard, and an example in OpenMetrics. https://github.com/open-telemetry/opentelemetry-specification/blob/89aa01348139/specification/metrics/data-model.md#exemplars https://github.com/OpenObservability/OpenMetrics/blob/138654493130/specification/OpenMetrics.md#exemplars-1 Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-02-15 14:20:08 +00:00
Bryan Boreham	12cac5bd5c	tsdb tests: use go-cmp instead of DeepEquals Also one simpler call checking nil. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-02-08 19:32:33 +00:00
Bryan Boreham	17f48f2b3b	Tests: use replacement DeepEquals in more places Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-02-08 19:32:33 +00:00
Bryan Boreham	39af788dbd	Tests: use replacement DeepEquals using go-cmp Use DeepEqual replacement using go-cmp, which is more flexible. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-02-08 19:30:20 +00:00
Peter Štibraný	e2b9cfeeeb	Enforce chunks ordering when writing index. (#8085 ) Document conditions on chunks. Add check on chunk time ordering. Signed-off-by: Peter Štibraný <peter.stibrany@grafana.com>	2024-02-04 16:31:49 +01:00
Bryan Boreham	98c4889029	Merge pull request #9298 from Creatone/creatone/use-testify tests: Move from t.Errorf and others.	2024-02-04 16:27:57 +01:00
Mikhail Fesenko	419dd265cc	Fix strange code, add messages to code brought in #8106 (#13509 ) Signed-off-by: Mikhail Fesenko <proggga@gmail.com>	2024-02-02 10:00:38 +01:00
Bryan Boreham	16e68c01e4	tests: remove err from message when testify prints it already For instance `require.NoError` will print the unexpected error; we don't need to include it in the message. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-02-01 14:18:01 +00:00
Mikhail Fesenko	5f2c3a5d3e	Small improvements, add const, remove copypasta (#8106 ) Signed-off-by: Mikhail Fesenko <proggga@gmail.com> Signed-off-by: Jesus Vazquez <jesusvzpg@gmail.com>	2024-02-01 14:30:50 +01:00
Paweł Szulik	5961f78186	Refactor tsdb tests to use testify. Signed-off-by: Paweł Szulik <paul.szulik@gmail.com>	2024-01-31 16:03:17 +00:00
Bryan Boreham	34230bb172	tsdb/wlog: close segment files sooner 'defer' runs at the end of the whole function; we should close each segment file as soon as we finished reading it. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-01-31 12:12:19 +00:00
Bryan Boreham	cd4562d3a6	Merge pull request #13473 from bboreham/pure-mutex tsdb: use cheaper Mutex on series	2024-01-30 09:57:08 +00:00
Marco Pracucci	501bc6419e	Add ShardedPostings() support to TSDB (#10421 ) This PR is a reference implementation of the proposal described in #10420. In addition to what described in #10420, in this PR I've introduced labels.StableHash(). The idea is to offer an hashing function which doesn't change over time, and that's used by query sharding in order to get a stable behaviour over time. The implementation of labels.StableHash() is the hashing function used by Prometheus before stringlabels, and what's used by Grafana Mimir for query sharding (because built before stringlabels was a thing). Follow up work As mentioned in #10420, if this PR is accepted I'm also open to upload another foundamental piece used by Grafana Mimir query sharding to accelerate the query execution: an optional, configurable and fast in-memory cache for the series hashes. Signed-off-by: Marco Pracucci <marco@pracucci.com>	2024-01-29 11:57:27 +00:00
Bryan Boreham	66237c1996	tsdb: use cheaper Mutex on series Mutex is 8 bytes; RWMutex is 24 bytes and much more complicated. Since `RLock` is only used in two places, `UpdateMetadata` and `Delete`, neither of which are hotspots, we should use the cheaper one. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-01-26 11:01:39 +00:00
Marco Pracucci	ec9cada56e	Remove unused isRegexMetaCharacter() Signed-off-by: Marco Pracucci <marco@pracucci.com>	2024-01-26 06:35:02 +01:00
Marco Pracucci	515890ec53	Use Matcher.SetMatches() Signed-off-by: Marco Pracucci <marco@pracucci.com>	2024-01-26 06:26:52 +01:00
Marco Pracucci	a1a45990a2	Fix TestPostingsForMatcher Signed-off-by: Marco Pracucci <marco@pracucci.com>	2024-01-25 14:59:39 +01:00
Bryan Boreham	b9eab6e4b8	tsdb: simplify internal series delete function (#13261 ) Lifting an optimisation from Agent code, `seriesHashmap.del` can use the unique series reference, doesn't need to check Labels. Also streamline the logic for deleting from `unique` and `conflicts` maps, and add some comments to help the next person. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2024-01-25 11:57:54 +01:00
Bryan Boreham	3f30ad3cc2	Merge pull request #13015 from bboreham/smaller-txring tsdb: make transaction isolation data structures smaller	2024-01-25 10:48:15 +00:00
Arve Knudsen	ba7012ec6a	TestHeadLabelValuesWithMatchers: Add test case (#13414 ) Add test case to TestHeadLabelValuesWithMatchers, while fixing a couple of typos in other test cases. Also enclosing some implicit sub-tests in a `t.Run` call to make them explicitly sub-tests. Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2024-01-24 10:47:56 +01:00
Filip Petkovski	583f3e587c	Optimize histogram iterators (#13340 ) Optimize histogram iterators Histogram iterators allocate new objects in the AtHistogram and AtFloatHistogram methods, which makes calculating rates over long ranges expensive. In #13215 we allowed an existing object to be reused when converting an integer histogram to a float histogram. This commit follows the same idea and allows injecting an existing object in the AtHistogram and AtFloatHistogram methods. When the injected value is nil, iterators allocate new histograms, otherwise they populate and return the injected object. The commit also adds a CopyTo method to Histogram and FloatHistogram which is used in the BufferedIterator to overwrite items in the ring instead of making new copies. Note that a specialized HPoint pool is needed for all of this to work (`matrixSelectorHPool`). --------- Signed-off-by: Filip Petkovski <filip.petkovsky@gmail.com> Co-authored-by: George Krajcsovits <krajorama@users.noreply.github.com>	2024-01-23 17:02:14 +01:00
Oleg Zaytsev	ed172a6667	Optimize label values with matchers by taking shortcuts (#13426 ) Don't calculate postings beforehand: we may not need them. If all matchers are for the requested label, we can just filter its values. Also, if there are no values at all, no need to run any kind of logic. Also add more labelValuesWithMatchers benchmarks Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2024-01-23 11:40:21 +01:00
Julien Pivotto	f52605b584	Merge pull request #13415 from aknuds1/arve/test-label-values-with-matchers-one-more TestLabelValuesWithMatchers: Add test case	2024-01-18 11:57:12 +01:00
tyltr	f97fa2736c	remove obsolete build tag Signed-off-by: tyltr <tylitianrui@126.com>	2024-01-17 22:26:32 +08:00
Arve Knudsen	8598150f48	TestLabelValuesWithMatchers: Add test case Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2024-01-17 12:51:56 +01:00
Marco Pracucci	7852a7c516	Fix regressions introduced by #13242 Signed-off-by: Marco Pracucci <marco@pracucci.com>	2024-01-16 12:00:53 +01:00
Giedrius Statkevičius	b695e069b8	tsdb/main: wire "EnableOverlappingCompaction" to tsdb.Options (#13398 ) This added the https://github.com/prometheus/prometheus/pull/13393 "EnableOverlappingCompaction" parameter to the compactor code but not to the tsdb.Options. I forgot about that. Add it to `tsdb.Options` too and set it to `true` in Prometheus. Copy/paste the description from https://github.com/prometheus/prometheus/pull/13393#issuecomment-1891787986 Signed-off-by: Giedrius Statkevičius <giedrius.statkevicius@vinted.com>	2024-01-15 16:42:40 +01:00
Ben Kochie	17920623e7	Merge pull request #13391 from GiedriusS/compact_merge_func tsdb/compact: fix passing merge func	2024-01-15 09:43:06 +01:00
Giedrius Statkevičius	3a48adc54f	tsdb: add enable overlapping compaction This functionality is needed in downstream projects because they have a separate component that does compaction. Upstreaming `7c8e9a2a76/tsdb/compact.go (L323-L325)`. Signed-off-by: Giedrius Statkevičius <giedrius.statkevicius@vinted.com>	2024-01-12 11:19:41 +02:00
Giedrius Statkevičius	9b759135d1	tsdb/compact: fix passing merge func Fixing a very small logical problem I've introduced :(. Signed-off-by: Giedrius Statkevičius <giedrius.statkevicius@vinted.com>	2024-01-11 12:07:54 +02:00
Giedrius Statkevičius	61b4080a14	tsdb/{index,compact}: allow using custom postings encoding format (#13242 ) * tsdb/{index,compact}: allow using custom postings encoding format We would like to experiment with a different postings encoding format in Thanos so in this change I am proposing adding another argument to `NewWriter` which would allow users to change the format if needed. Also, wire the leveled compactor so that it would be possible to change the format there too. Signed-off-by: Giedrius Statkevičius <giedrius.statkevicius@vinted.com> * tsdb/compact: use a struct for leveled compactor options As discussed on Slack, let's use a struct for the options in leveled compactor. Signed-off-by: Giedrius Statkevičius <giedrius.statkevicius@vinted.com> * tsdb: make changes after Bryan's review - Make changes less intrusive - Turn the postings encoder type into a function - Add NewWriterWithEncoder() Signed-off-by: Giedrius Statkevičius <giedrius.statkevicius@vinted.com> --------- Signed-off-by: Giedrius Statkevičius <giedrius.statkevicius@vinted.com>	2024-01-08 09:48:27 +00:00
Bryan Boreham	bad3f23f23	agent: add BenchmarkCreateSeries Based on the one in tsdb/head_test.go. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2023-12-31 10:23:43 +00:00
Bryan Boreham	e64d7d8928	agent: make the global hash lookup table smaller This is the same change made in #13040, plus subsequent improvements, applied to agent-mode code. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2023-12-31 10:23:43 +00:00
Bryan Boreham	252031c86f	Revert "Adding small test update for temp dir using t.TempDir (#13293 )" This reverts commit `2ddb3596ef`. Various tests are failing in CI after this change; reverting to free up other work. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2023-12-30 19:17:30 +00:00
Mile Druzijanic	2ddb3596ef	Adding small test update for temp dir using t.TempDir (#13293 ) * Adding small test update for temp dir using t.TempDir Signed-off-by: Mile Druzijanic <miledruz@gmail.com> Signed-off-by: Mile Druzijanic <zedsprogramms@gmail.com> * removing not required cleanup Signed-off-by: Mile Druzijanic <zedsprogramms@gmail.com> --------- Signed-off-by: Mile Druzijanic <miledruz@gmail.com> Signed-off-by: Mile Druzijanic <zedsprogramms@gmail.com>	2023-12-28 21:49:57 +01:00
Björn Rabenstein	6b8e945388	Merge pull request #13289 from fpetkovski/fix-histogram-reuse Fix reusing float histograms	2023-12-25 22:45:03 +01:00
Bryan Boreham	8065bef172	Move metric type definitions to common/model They are used in multiple repos, so common is a better place for them. Several packages now don't depend on `model/textparse`, e.g. `storage/remote`. Also remove `metadata` struct from `api.go`, since it was identical to a struct in the `metadata` package. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2023-12-19 18:56:54 +00:00
Filip Petkovski	1f69dcfa6b	Fix reusing float histograms In https://github.com/prometheus/prometheus/pull/13276 we started reusing float histogram objects to reduce allocations in PromQL. That PR introduces a bug where histogram pointers gets copied to the beginning of the histograms slice, but are still kept in the end of the slice. When a new histogram is read into the last element, it can overwrite a previous element because the pointer is the same. This commit fixes the issue by moving outdated points to the end of the slice so that we don't end up with duplicate pointers in the same buffer. In other words, the slice gets rotated so that old objects can get reused. Signed-off-by: Filip Petkovski <filip.petkovsky@gmail.com>	2023-12-14 11:53:58 +01:00
Bryan Boreham	d0c2d9c0b9	Merge pull request #12878 from bboreham/loser-tree postings: use Loser Tree for merge	2023-12-12 21:38:30 +00:00
Björn Rabenstein	928d07e3bd	Merge branch 'main' into arve/typos Signed-off-by: Björn Rabenstein <beorn@grafana.com>	2023-12-12 12:02:03 +01:00
Giedrius Statkevičius	f36b56a62c	tsdb: remove unused option (#13282 ) Digging around the TSDB code and I've found that this flag is unused so let's remove it. Signed-off-by: Giedrius Statkevičius <giedrius.statkevicius@vinted.com>	2023-12-12 09:58:54 +00:00

1 2 3 4 5 ...

1167 commits