prometheus

mirror of https://github.com/prometheus/prometheus.git synced 2024-12-28 23:19:41 -08:00

Author	SHA1	Message	Date
Goutham Veeramachaneni	86729d4d7b	Update exp package (#12650 )	2023-09-21 22:53:51 +02:00
Björn Rabenstein	f8dd8770ac	Merge pull request #12757 from bboreham/reuse-bufiter TSDB: re-use iterator when moving between series	2023-09-21 14:08:53 +02:00
Dimitar Dimitrov	1155d736b6	Improve sensitivity of TestQuerierIndexQueriesRace Currently, the two goroutines race against each other and it's possible that the main test goroutine finishes way earlier than appendSeries has had a chance to run at all. I tested this change by breaking the code that X fixed and running the race test 100 times. Without the additional time.Sleep the test failed 11 times. With the sleep it failed 65 out of the 100 runs. Which is still not ideal, but it's a step forward. Signed-off-by: Dimitar Dimitrov <dimitar.dimitrov@grafana.com>	2023-09-21 12:30:08 +02:00
Dimitar Dimitrov	6f1284ac93	Fix exit condition of TestQuerierIndexQueriesRace The test was introduced in # but was changed during the code review and not reran with the faulty code since then. Closes # Signed-off-by: Dimitar Dimitrov <dimitar.dimitrov@grafana.com>	2023-09-20 20:22:26 +01:00
Björn Rabenstein	864da019cd	Merge pull request #12874 from krajorama/outof-order-chunks Fix duplicate sample detection at chunk size limit	2023-09-20 18:01:21 +02:00
Björn Rabenstein	9071913fd9	Merge pull request #12831 from aknuds1/arve/posting-context Add context argument to `tsdb.PostingsForMatchers`	2023-09-20 17:15:15 +02:00
György Krajcsovits	9dbd100a5e	Refactor solution to not repeat code Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com>	2023-09-20 15:54:00 +02:00
György Krajcsovits	96d03b6f46	Fix duplicate sample detection at chunks size limit Before cutting a new XOR chunk in case the chunk goes over the size limit, check that the timestamp is in order and not equal or older than the latest sample in the old chunk. Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com>	2023-09-20 14:49:56 +02:00
György Krajcsovits	56b3a015b6	Add regression test for duplicate detection at chunk size limit TestHeadDetectsDuplcateSampleAtSizeLimit tests a regression where a duplicate sample,is appended to the head, right when the head chunk is at the size limit. The test adds all samples as duplicate, thus expecting that the result has exactly half of the samples. Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com>	2023-09-20 14:32:20 +02:00
Björn Rabenstein	83891135c6	Merge pull request #12838 from krajorama/fix-disappearing-span-panic Fix counterResetInAnyBucket panic	2023-09-19 17:10:27 +02:00
George Krajcsovits	3512b2d678	storage: make histogram reset handling consistent in chainSampleIterator (#12779 ) storage: make histogram reset handling consistent in chainSampleIterator --------- Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com>	2023-09-19 17:06:46 +02:00
Alan Protasio	959c98441b	Add context argument to tsdb.PostingsForMatchers Signed-off-by: Alan Protasio <alanprot@gmail.com>	2023-09-16 18:13:32 +02:00
zenador	69edd8709b	Add warnings (and annotations) to PromQL query results (#12152 ) Return annotations (warnings and infos) from PromQL queries This generalizes the warnings we have already used before (but only for problems with remote read) as "annotations". Annotations can be warnings or infos (the latter could be false positives). We do not treat them different in the API for now and return them all as "warnings". It would be easy to distinguish them and return infos separately, should that appear useful in the future. The new annotations are then used to create a lot of warnings or infos during PromQL evaluations. Partially these are things we have wanted for a long time (e.g. inform the user that they have applied `rate` to a metric that doesn't look like a counter), but the new native histograms have created even more needs for those annotations (e.g. if a query tries to aggregate float numbers with histograms). The annotations added here are not yet complete. A prominent example would be a warning about a range too short for a rate calculation. But such a warnings is more tricky to create with good fidelity and we will tackle it later. Another TODO is to take annotations into account when evaluating recording rules. --------- Signed-off-by: Jeanette Tan <jeanette.tan@grafana.com>	2023-09-14 18:57:31 +02:00
Arve Knudsen	156222cc50	Add context argument to LabelQuerier.LabelValues (#12665 ) Add context argument to LabelQuerier.LabelValues and LabelQuerier.SortedLabelValues. Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2023-09-14 16:02:04 +02:00
Arve Knudsen	a964349e97	Add context argument to LabelQuerier.LabelNames (#12666 ) Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2023-09-14 10:39:51 +02:00
Arve Knudsen	4451ba10b4	Add context argument to IndexReader.Postings (#12667 ) Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2023-09-13 17:45:06 +02:00
Arve Knudsen	6ef9ed0bc3	Add context argument to DB.Delete (#12834 ) Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2023-09-13 15:43:06 +02:00
György Krajcsovits	b2fa4d910a	Fix more counterResetInAnyBucket edgecases Case a) empty span is at the beginning of the spans. Case b) two consequtive empty spans with positive offsets. Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com>	2023-09-13 15:18:40 +02:00
Fiona Liao	4419399e4e	Do WBL mmap marker replay concurrently (#12801 ) * Benchmark WBL Extended WAL benchmark test with WBL parts too - added basic cases for OOO handling - a percentage of series have a percentage of samples set as OOO ones. Signed-off-by: Fiona Liao <fiona.y.liao@gmail.com>	2023-09-12 21:31:10 +02:00
Shirley	d3a1044354	WBL loading: don't send empty buffers over chan (#12808 ) Signed-off-by: Shirley Leu <4163034+fridgepoet@users.noreply.github.com> Co-authored-by: Fiona Liao <fiona.y.liao@gmail.com>	2023-09-12 16:26:02 +02:00
Arve Knudsen	6daee89e5f	Add context argument to Querier.Select (#12660 ) Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2023-09-12 12:37:38 +02:00
Bryan Boreham	f711d71aa8	Merge pull request #12798 from fionaliao/remove-duplicated-max-time Remove duplicated ms.mmMaxTime check in processWALSamples	2023-09-06 09:17:42 +01:00
Fiona Liao	f211fcd92d	Remove duplicated ms.mmMaxTime check in WAL Signed-off-by: Fiona Liao <fiona.y.liao@gmail.com>	2023-09-05 15:23:03 +01:00
George Krajcsovits	b6f903b5f9	Fix handling of explicit counter reset header in histograms. (#12772 ) * Fix handling of explicit counter reset header in histograms. Explicit counter reset were being ignored. Also there was no unit test coverage. Add test case for the first sample in a chunk. Add test case for non first sample in chunk. Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com> --------- Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com>	2023-09-01 23:39:15 +02:00
Dimitar Dimitrov	b40865833d	PostingsForMatchers race with creating new series (#12558 ) Signed-off-by: Dimitar Dimitrov <dimitar.dimitrov@grafana.com>	2023-08-29 11:03:27 +02:00
Bryan Boreham	c5671c6d97	Merge pull request #12755 from bboreham/rangequery-benchmark-mmap promql: force mmap of head chunks in BenchmarkRangeQuery	2023-08-26 15:56:52 +01:00
Bryan Boreham	bdc7983956	TSDB: re-use iterator when moving between series Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2023-08-26 14:01:44 +00:00
Bryan Boreham	5d22d422ab	Merge pull request #12690 from michalbiesek/feat-go-bump Update Go version to 1.21	2023-08-26 14:36:15 +01:00
Bryan Boreham	0d283effa8	promql: force mmap of head chunks in BenchmarkRangeQuery Otherwise we have a highly unusual situation of over 100 chunks in the headChunks list of each series, which heavily skews performance. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2023-08-26 09:40:59 +00:00
Gregor Zeitlinger	f01718262a	Unit tests for native histograms (#12668 ) promql: Extend testing framework to support native histograms This includes both the internal testing framework as well as the rules unit test feature of promtool. This also adds a bunch of basic tests. Many of the code level tests can now be converted to tests within the framework, and more tests can be added easily. --------- Signed-off-by: Harold Dost <h.dost@criteo.com> Signed-off-by: Gregor Zeitlinger <gregor.zeitlinger@grafana.com> Signed-off-by: Stephen Lang <stephen.lang@grafana.com> Co-authored-by: Harold Dost <h.dost@criteo.com> Co-authored-by: Stephen Lang <stephen.lang@grafana.com> Co-authored-by: Gregor Zeitlinger <gregor.zeitlinger@grafana.com>	2023-08-25 23:35:42 +02:00
Michal Biesek	04d7b4dbee	lint: Fix `SA1019` Using a deprecated function `rand.Read` has been deprecated since Go 1.20 `crypto/rand.Read` is more appropriate Ref: https://tip.golang.org/doc/go1.20 Signed-off-by: Michal Biesek <michalbiesek@gmail.com>	2023-08-25 17:47:41 +02:00
Justin Lei	8ef7dfdeeb	Add a chunk size limit in bytes (#12054 ) Add a chunk size limit in bytes This creates a hard cap for XOR chunks of 1024 bytes. The limit for histogram chunk is also 1024 bytes, but it is a soft limit as a histogram has a dynamic size, and even a single one could be larger than 1024 bytes. This also avoids cutting new histogram chunks if the existing chunk has fewer than 10 histograms yet. In that way, we are accepting "jumbo chunks" in order to have at least 10 histograms in a chunk, allowing compression to kick in. Signed-off-by: Justin Lei <justin.lei@grafana.com>	2023-08-24 15:21:17 +02:00
beorn7	aa82fe198f	tsdb: Fix histogram validation So far, `ValidateHistogram` would not detect if the count did not include the count in the zero bucket. This commit fixes the problem and updates all the tests that have been undetected offenders so far. Note that this problem would only ever create false negatives, so we never falsely rejected to store a histogram because of it. On the other hand, `ValidateFloatHistogram` has been to strict with the count being at least as large as the sum of the counts in all the buckets. Float precision issues could create false positives here, see products of PromQL evaluations, it's actually quite hard to put an upper limit no the floating point imprecision. Users could produce the weirdest expressions, maxing out float precision problems. Therefore, this commit simply removes that particular check from `ValidateFloatHistogram`. Signed-off-by: beorn7 <beorn@grafana.com>	2023-08-22 23:04:01 +02:00
Mustafa Ateş Uzun	e5e51bebef	fix: error message typo Signed-off-by: Mustafa Ateş Uzun <mustafauzun0@gmail.com>	2023-08-17 16:34:45 +03:00
Julien Pivotto	e3fabd5fdf	Merge pull request #12664 from prometheus/superq/cleanup_chunk_snapshots Cleanup temporary chunk snapshot dirs	2023-08-08 13:02:39 +02:00
SuperQ	8d38d59fc5	Cleanup temporary chunk snapshot dirs Simlar to cleanup of WAL files on startup, cleanup temporary chunk_snapshot dirs. This prevents storage space leaks due to terminated snapshots on shutdown. Signed-off-by: SuperQ <superq@gmail.com>	2023-08-08 09:43:48 +02:00
Julien Pivotto	c3311272d9	Merge pull request #12652 from colega/fix-typo-in-append-histogram-param-name Fix typo in Appender.AppendHistogram() arg name	2023-08-04 16:37:40 +02:00
Oleg Zaytsev	6ea6def0d3	Use zeropool when replaying agent's DB WAL (#12651 ) Same as https://github.com/prometheus/prometheus/pull/12189 but for tsdb/agent/db.go Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2023-08-04 10:39:55 +02:00
Oleg Zaytsev	c810e7cae3	Fix typo in Appender.AppendHistogram() arg name Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2023-08-04 10:21:16 +02:00
Oleg Zaytsev	61daa30bb1	Pass ref to SeriesLifecycleCallback.PostDeletion (#12626 ) When a particular SeriesLifecycleCallback tries to optimize and run closer to the Head, keeping track of the HeadSeriesRef instead of the labelsets, it's impossible to handle the PostDeletion callback properly as there's no way to know which series refs were deleted from the head. This changes the callback to provide the series refs alongside the labelsets, so the implementation can choose what to do. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2023-08-03 10:56:27 +02:00
Oleg Zaytsev	cd7d0b69a2	Check nil err first when committing (#12625 ) The most common case is to have a nil error when appending series, so let's check that first instead of checking the 3 error types first. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2023-08-01 14:04:45 +02:00
cui fliter	f26dfc95e6	fix struct name in comment (#12624 ) Signed-off-by: cui fliter <imcusg@gmail.com>	2023-08-01 12:24:42 +02:00
Łukasz Mierzwa	3c80963e81	Use a linked list for memSeries.headChunk (#11818 ) Currently memSeries holds a single head chunk in-memory and a slice of mmapped chunks. When append() is called on memSeries it might decide that a new headChunk is needed to use for given append() call. If that happens it will first mmap existing head chunk and only after that happens it will create a new empty headChunk and continue appending our sample to it. Since appending samples uses write lock on memSeries no other read or write can happen until any append is completed. When we have an append() that must create a new head chunk the whole memSeries is blocked until mmapping of existing head chunk finishes. Mmapping itself uses a lock as it needs to be serialised, which means that the more chunks to mmap we have the longer each chunk might wait for it to be mmapped. If there's enough chunks that require mmapping some memSeries will be locked for long enough that it will start affecting queries and scrapes. Queries might timeout, since by default they have a 2 minute timeout set. Scrapes will be blocked inside append() call, which means there will be a gap between samples. This will first affect range queries or calls using rate() and such, since the time range requested in the query might have too few samples to calculate anything. To avoid this we need to remove mmapping from append path, since mmapping is blocking. But this means that when we cut a new head chunk we need to keep the old one around, so we can mmap it later. This change makes memSeries.headChunk a linked list, memSeries.headChunk still points to the 'open' head chunk that receives new samples, while older, yet to be mmapped, chunks are linked to it. Mmapping is done on a schedule by iterating all memSeries one by one. Thanks to this we control when mmapping is done, since we trigger it manually, which reduces the risk that it will have to compete for mmap locks with other chunks. Signed-off-by: Łukasz Mierzwa <l.mierzwa@gmail.com>	2023-07-31 11:10:24 +02:00
Robert Fratto	886945cda7	tsdb/agent: ensure that new series get written to WAL on rollback (#12592 ) If a new series is introduced in a storage.Appender instance, that series should be written to the WAL once the storage.Appender is closed, even on Rollback. Previously, new series would only be written to the WAL when calling Commit. However, because the series is stored in memory regardless, subsequent calls to Commit may write samples to the WAL which reference a series ID which that was never written. Related to #11589. It's likely that this fix also resolves this issue, but we need more testing from users to see if the problem persists after this fix; there may be more cases where samples get written to the WAL in Prometheus Agent mode without the corresponding series record. Signed-off-by: Robert Fratto <robertfratto@gmail.com>	2023-07-27 09:28:26 -04:00
George Krajcsovits	6cd2d1621f	Hide histogram chunk append and reset header internals (#12352 ) tsdb: Hide histogram chunk append and reset header internals Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com> Signed-off-by: George Krajcsovits <krajorama@users.noreply.github.com>	2023-07-26 15:08:16 +02:00
Björn Rabenstein	0e12f11d61	Merge pull request #12583 from prometheus/release-2.46 Merge release-2.46 into main	2023-07-20 18:29:44 +02:00
György Krajcsovits	d4e355243a	tsdbutil/ChunkFromSamplesGeneric should not panic Add error handling instead. Prepares for #12352 Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com>	2023-07-20 17:01:34 +02:00
Julien Pivotto	7905594b52	Merge pull request #12557 from prometheus/beorn7/histogram scrape: Enable ingestion of multiple exemplars per sample	2023-07-20 15:19:28 +02:00
Julien Pivotto	1f5934e7be	Merge pull request #10623 from songjiayang/update-index make sure response error when TOC parse failed	2023-07-18 13:47:27 +02:00
cui fliter	096ceca44f	remove repetitive words (#12556 ) Signed-off-by: cui fliter <imcusg@gmail.com>	2023-07-13 15:53:40 +02:00
beorn7	0e3f35324b	scrape: Enable ingestion of multiple exemplars per sample This has become a requirement for native histograms, as a single histogram sample commonly has many buckets, so that providing many exemplars makes sense. Since OM text doesn't support native histograms yet, the test had to be expanded to also support protobuf test cases. Signed-off-by: beorn7 <beorn@grafana.com>	2023-07-13 14:16:10 +02:00
Julien Pivotto	89e213bc02	Merge pull request #12546 from roidelapluie/removeimport TSDB: Remove usused import of sort	2023-07-11 15:06:48 +02:00
Justin Lei	32d87282ad	Add Zstandard compression option for wlog (#11666 ) Snappy remains as the default compression but there is now a flag to switch the compression algorithm. Signed-off-by: Justin Lei <justin.lei@grafana.com>	2023-07-11 14:57:57 +02:00
Julien Pivotto	bf5bf1a4b3	TSDB: Remove usused import of sort Signed-off-by: Julien Pivotto <roidelapluie@o11y.eu>	2023-07-11 14:29:31 +02:00
Julien Pivotto	8c8afec116	Merge pull request #12542 from merrickclay/tsdb-doc-comment improve incorrect doc comment	2023-07-11 13:10:04 +02:00
Julien Pivotto	0f85e4f41d	Merge pull request #12539 from bboreham/slices-sorts Replace sort.Slice with faster slices.SortFunc	2023-07-11 13:09:02 +02:00
Merrick Clay	70e41fc5ac	improve incorrect doc comment Signed-off-by: Merrick Clay <merrick.e.clay@gmail.com>	2023-07-10 16:52:00 -06:00
Bryan Boreham	ce153e3fff	Replace sort.Sort with faster slices.SortFunc The generic version is more efficient. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2023-07-10 09:43:45 +00:00
Marc Tudurí	4851ced266	tsdb: Support native histograms in snapshot on shutdown (#12258 ) Signed-off-by: Marc Tuduri <marctc@protonmail.com>	2023-07-05 11:44:13 +02:00
Julien Pivotto	9ff1f24efa	Merge pull request #12505 from pracucci/fix-infinite-loop-in-index-writer Fix infinite loop in index Writer when a series contains duplicated label names	2023-07-04 13:08:36 +02:00
Patrick Oyarzun	68e5937474	Apply relevant label matchers in LabelValues before fetching extra postings (#12274 ) * Apply matchers when fetching label values Signed-off-by: Patrick Oyarzun <patrick.oyarzun@grafana.com> * Avoid extra copying of label values Signed-off-by: Patrick Oyarzun <patrick.oyarzun@grafana.com> --------- Signed-off-by: Patrick Oyarzun <patrick.oyarzun@grafana.com>	2023-07-04 10:37:58 +01:00
Bryan Boreham	5255bf06ad	Replace sort.Slice with faster slices.SortFunc The generic version is more efficient. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2023-07-02 22:17:08 +00:00
Marco Pracucci	35069910f5	Fix infinite loop in index Writer when a series contains duplicated label names Signed-off-by: Marco Pracucci <marco@pracucci.com>	2023-07-01 17:38:08 +02:00
Marco Pracucci	031d22df9e	Fix race condition in ChunkDiskMapper.Truncate() (#12500 ) * Fix race condition in ChunkDiskMapper.Truncate() Signed-off-by: Marco Pracucci <marco@pracucci.com> * Added unit test Signed-off-by: Marco Pracucci <marco@pracucci.com> * Update tsdb/chunks/head_chunks.go Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com> Signed-off-by: Marco Pracucci <marco@pracucci.com> --------- Signed-off-by: Marco Pracucci <marco@pracucci.com> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-06-30 18:29:59 +05:30
Bartlomiej Plotka	4062f12573	Merge pull request #12396 from leizor/leizor/chunk-opts Group args to append to memSeries in chunkOpts	2023-06-27 13:08:21 +02:00
Nidhey Nitin Indurkar	a8772a4178	Feat: Get block by id directly on promtool analyze & get latest block if ID not provided (#12031 ) * feat: analyze latest block or block by ID in CLI (promtool) Signed-off-by: nidhey27 <nidhey.indurkar@infracloud.io> * address remarks Signed-off-by: nidhey60@gmail.com <nidhey.indurkar@infracloud.io> * address latest review comments Signed-off-by: nidhey60@gmail.com <nidhey.indurkar@infracloud.io> --------- Signed-off-by: nidhey27 <nidhey.indurkar@infracloud.io> Signed-off-by: nidhey60@gmail.com <nidhey.indurkar@infracloud.io>	2023-06-01 17:13:09 +05:30
Alan Protasio	73078bf738	Opmizing Group Regex (#12375 ) Signed-off-by: Alan Protasio <alanprot@gmail.com>	2023-05-30 13:49:22 +02:00
Julien Pivotto	6f97641a51	Merge pull request #12380 from mmorel-35/patch-2 ci(lint): enable predeclared linter	2023-05-28 14:43:29 +02:00
Justin Lei	e73d8b2084	Also pass chunkOpts into appendPreprocessor Signed-off-by: Justin Lei <justin.lei@grafana.com>	2023-05-25 13:37:18 -07:00
Justin Lei	4c4454e4c9	Group args to append to memSeries in chunkOpts Signed-off-by: Justin Lei <justin.lei@grafana.com>	2023-05-25 13:12:46 -07:00
Justin Lei	89af351730	Remove samplesPerChunk from memSeries (#12390 ) Signed-off-by: Justin Lei <justin.lei@grafana.com>	2023-05-25 11:18:41 +02:00
zenador	37e5249e33	Use DefaultSamplesPerChunk in tsdb (#12387 ) Signed-off-by: Jeanette Tan <jeanette.tan@grafana.com>	2023-05-24 13:00:21 +02:00
Baskar Shanmugam	905a0bd63a	Added 'limit' query parameter support to /api/v1/status/tsdb endpoint (#12336 ) * Added 'topN' query parameter support to /api/v1/status/tsdb endpoint Signed-off-by: Baskar Shanmugam <baskar.shanmugam.career@gmail.com> * Updated query parameter for tsdb status to 'limit' Signed-off-by: Baskar Shanmugam <baskar.shanmugam.career@gmail.com> * Corrected Stats() parameter name from topN to limit Signed-off-by: Baskar Shanmugam <baskar.shanmugam.career@gmail.com> * Fixed p.Stats CI failure Signed-off-by: Baskar Shanmugam <baskar.shanmugam.career@gmail.com> --------- Signed-off-by: Baskar Shanmugam <baskar.shanmugam.career@gmail.com>	2023-05-22 14:37:07 +02:00
Alan Protasio	8c5d4b4add	Opmize MatchNotEqual (#12377 ) Signed-off-by: Alan Protasio <alanprot@gmail.com>	2023-05-21 10:41:30 +02:00
Matthieu MOREL	c8e7f95a3c	ci(lint): enable predeclared linter Signed-off-by: Matthieu MOREL <matthieu.morel35@gmail.com>	2023-05-21 07:33:54 +00:00
George Krajcsovits	92d6980360	Fix populateWithDelChunkSeriesIterator and gauge histograms (#12330 ) Use AppendableGauge to detect corrupt chunk with gauge histograms. Detect if first sample is a gauge but the chunk is not set up to contain gauge histograms. Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com> Signed-off-by: George Krajcsovits <krajorama@users.noreply.github.com>	2023-05-19 10:24:06 +02:00
Baskar Shanmugam	f731a90a7f	Fix LabelValueStats in posting stats (#12342 ) Problem: LabelValueStats - This will provide a list of the label names and memory used in bytes. It is calculated by adding the length of all values for a given label name. But internally Prometheus stores the name and the value independently for each series. Solution: MemPostings struct maintains the values to seriesRef map which is used to get the number of series which contains the label values. Using that LabelValueStats is calculated as: seriesCnt * len(value name) Signed-off-by: Baskar Shanmugam <baskar.shanmugam.career@gmail.com>	2023-05-19 09:36:30 +02:00
Xiaochao Dong	80b7f73d26	Copy tombstone intervals to avoid race (#12245 ) Signed-off-by: Xiaochao Dong (@damnever) <the.xcdong@gmail.com>	2023-05-17 15:15:12 +02:00
Björn Rabenstein	30e263cf96	Merge pull request #12357 from krajorama/fix-histogram-appendable-emptybucket fix HistogramAppender.appendable segfault	2023-05-16 20:52:39 +02:00
Callum Styan	0d2108ad79	[tsdb] re-implement WAL watcher to read via a "notification" channel (#11949 ) * WIP implement WAL watcher reading via notifications over a channel from the TSDB code Signed-off-by: Callum Styan <callumstyan@gmail.com> * Notify via head appenders Commit (finished all WAL logging) rather than on each WAL Log call Signed-off-by: Callum Styan <callumstyan@gmail.com> * Fix misspelled Notify plus add a metric for dropped Write notifications Signed-off-by: Callum Styan <callumstyan@gmail.com> * Update tests to handle new notification pattern Signed-off-by: Callum Styan <callumstyan@gmail.com> * this test maybe needs more time on windows? Signed-off-by: Callum Styan <callumstyan@gmail.com> * does this test need more time on windows as well? Signed-off-by: Callum Styan <callumstyan@gmail.com> * read timeout is already a time.Duration Signed-off-by: Callum Styan <callumstyan@gmail.com> * remove mistakenly commited benchmark data files Signed-off-by: Callum Styan <callumstyan@gmail.com> * address some review feedback Signed-off-by: Callum Styan <callumstyan@gmail.com> * fix missed changes from previous commit Signed-off-by: Callum Styan <callumstyan@gmail.com> * Fix issues from wrapper function Signed-off-by: Callum Styan <callumstyan@gmail.com> * try fixing race condition in test by allowing tests to overwrite the read ticker timeout instead of calling the Notify function Signed-off-by: Callum Styan <callumstyan@gmail.com> * fix linting Signed-off-by: Callum Styan <callumstyan@gmail.com> --------- Signed-off-by: Callum Styan <callumstyan@gmail.com>	2023-05-15 12:31:49 -07:00
György Krajcsovits	c6618729c9	Fix HistogramAppender.Appendable array out of bound error The code did not handle spans with 0 length properly. Spans with length zero are now skipped in the comparison. Span index check not done against length-1, since length is a unit32, thus subtracting 1 leads to 2^32, not -1. Fixes and unit tests for both integer and float histograms added. Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com>	2023-05-14 17:38:52 +02:00
Jesus Vazquez	1f1dac2cda	Merge pull request #12351 from alanprot/optimization/MatchNotRegexp Implementing Regex optimization on the `MatchNotRegexp` matcher type	2023-05-11 11:55:17 +02:00
Alan Protasio	c0f1abb574	MatchNotRegexp optimization Signed-off-by: Alan Protasio <alanprot@gmail.com>	2023-05-10 20:08:38 -07:00
Robert Fratto	9e4e2a4a51	wlog: use filepath for getting checkpoint number This changes usage of path to be replaced with path/filepath, allowing for filepath.Base to properly return the base directory on systems where `/` is not the standard path separator. This resolves an issue on Windows where intermediate folders containing a `.` were incorrectly considered to be a part of the checkpoint name. Related to grafana/agent#3826. Signed-off-by: Robert Fratto <robertfratto@gmail.com>	2023-05-10 12:38:02 -04:00
Björn Rabenstein	37fe9b89dc	Merge pull request #12055 from leizor/leizor/prometheus/issues/12009 Adjust samplesPerChunk from 120 to 220	2023-05-10 14:45:12 +02:00
Bryan Boreham	0ab9553611	tsdb: drop deleted series from the WAL sooner (#12297 ) `head.deleted` holds the WAL segment in use at the time each series was removed from the head. At the end of `truncateWAL()` we will delete all segments up to `last`, so we can drop any series that were last seen in a segment at or before that point. (same change in Prometheus Agent too) Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2023-05-01 16:43:15 +01:00
cui fliter	276ca6a883	fix some comments Signed-off-by: cui fliter <imcusg@gmail.com>	2023-04-25 14:19:16 +08:00
Matthieu MOREL	bae9a21200	Merge branch 'main' into linter/nilerr Signed-off-by: Matthieu MOREL <matthieu.morel35@gmail.com>	2023-04-19 19:56:39 +02:00
beorn7	5b53aa1108	style: Replace `else if` cascades with `switch` Wiser coders than myself have come to the conclusion that a `switch` statement is almost always superior to a statement that includes any `else if`. The exceptions that I have found in our codebase are just these two: * The `if else` is followed by an additional statement before the next condition (separated by a `;`). * The whole thing is within a `for` loop and `break` statements are used. In this case, using `switch` would require tagging the `for` loop, which probably tips the balance. Why are `switch` statements more readable? For one, fewer curly braces. But more importantly, the conditions all have the same alignment, so the whole thing follows the natural flow of going down a list of conditions. With `else if`, in contrast, all conditions but the first are "hidden" behind `} else if `, harder to spot and (for no good reason) presented differently from the first condition. I'm sure the aforemention wise coders can list even more reasons. In any case, I like it so much that I have found myself recommending it in code reviews. I would like to make it a habit in our code base, without making it a hard requirement that we would test on the CI. But for that, there has to be a role model, so this commit eliminates all `if else` occurrences, unless it is autogenerated code or fits one of the exceptions above. Signed-off-by: beorn7 <beorn@grafana.com>	2023-04-19 17:22:31 +02:00
beorn7	c3c7d44d84	lint: Adjust to the lint warnings raised by current versions of golint-ci We haven't updated golint-ci in our CI yet, but this commit prepares for that. There are a lot of new warnings, and it is mostly because the "revive" linter got updated. I agree with most of the new warnings, mostly around not naming unused function parameters (although it is justified in some cases for documentation purposes – while things like mocks are a good example where not naming the parameter is clearer). I'm pretty upset about the "empty block" warning to include `for` loops. It's such a common pattern to do something in the head of the `for` loop and then have an empty block. There is still an open issue about this: https://github.com/mgechev/revive/issues/810 I have disabled "revive" altogether in files where empty blocks are used excessively, and I have made the effort to add individual `// nolint:revive` where empty blocks are used just once or twice. It's borderline noisy, though, but let's go with it for now. I should mention that none of the "empty block" warnings for `for` loop bodies were legitimate. Signed-off-by: beorn7 <beorn@grafana.com>	2023-04-19 17:10:10 +02:00
Đurica Yuri Nikolić	b028112331	Making the number of CPU cores used for sorting postings lists editable (#12247 ) Signed-off-by: Yuri Nikolic <durica.nikolic@grafana.com>	2023-04-18 12:13:05 +02:00
Ganesh Vernekar	7309ac2721	Merge pull request #12257 from alexqyle/block-populator-rename Rename PopulateBlockFunc to BlockPopulator	2023-04-14 13:35:01 +08:00
Justin Lei	c3e6b85631	Reverse test changes Signed-off-by: Justin Lei <justin.lei@grafana.com>	2023-04-13 15:59:49 -07:00
Justin Lei	052993414a	Add storage.tsdb.samples-per-chunk flag Signed-off-by: Justin Lei <justin.lei@grafana.com>	2023-04-13 15:59:49 -07:00
Matthieu MOREL	fb3eb21230	enable gocritic, unconvert and unused linters Signed-off-by: Matthieu MOREL <matthieu.morel35@gmail.com>	2023-04-13 19:20:22 +00:00
beorn7	817a2396cb	Name float values as "floats", not as "values" In the past, every sample value was a float, so it was fine to call a variable holding such a float "value" or "sample". With native histograms, a sample might have a histogram value. And a histogram value is still a value. Calling a float value just "value" or "sample" or "V" is therefore misleading. Over the last few commits, I already renamed many variables, but this cleans up a few more places where the changes are more invasive. Note that we do not to attempt naming in the JSON APIs or in the protobufs. That would be quite a disruption. However, internally, we can call variables as we want, and we should go with the option of avoiding misunderstandings. Signed-off-by: beorn7 <beorn@grafana.com>	2023-04-13 19:25:24 +02:00
beorn7	630bcb494b	storage: Use separate sample types for histogram vs. float Previously, we had one “polymorphous” `sample` type in the `storage` package. This commit breaks it up into `fSample`, `hSample`, and `fhSample`, each still implementing the `tsdbutil.Sample` interface. This reduces allocations in `sampleRing.Add` but inflicts the penalty of the interface wrapper, which makes things worse in total. This commit therefore just demonstrates the step taken. The next commit will tackle the interface overhead problem. Signed-off-by: beorn7 <beorn@grafana.com>	2023-04-13 19:25:24 +02:00
Alex Le	01d0dda4fc	Rename PopulateBlockFunc to BlockPopulator Signed-off-by: Alex Le <leqiyue@amazon.com>	2023-04-12 14:18:20 -07:00
Björn Rabenstein	8ed90b567b	Merge pull request #12234 from aknuds1/chore/improve-histogram-comments tsdb: Improve a couple of histogram documentation comments	2023-04-12 10:55:22 +02:00
Björn Rabenstein	6e0a46900b	Merge pull request #12192 from leizor/leizor/prometheus/issues/11204 Add support for native histograms to concreteSeriesIterator	2023-04-11 12:30:35 +02:00
Arve Knudsen	cca7178a12	tsdb: Improve a couple of histogram documentation comments Signed-off-by: Arve Knudsen <arve.knudsen@gmail.com>	2023-04-07 18:06:27 +02:00
Justin Lei	83f43982c9	Add support for native histograms to concreteSeriesIterator Signed-off-by: Justin Lei <justin.lei@grafana.com>	2023-04-06 09:54:15 -07:00
Justin Lei	73ff91d182	Test fixes Signed-off-by: Justin Lei <justin.lei@grafana.com>	2023-04-06 09:42:59 -07:00
Justin Lei	c770ba8047	Add comment linking to PR Signed-off-by: Justin Lei <justin.lei@grafana.com>	2023-04-06 09:19:32 -07:00
Justin Lei	79db04eb12	Adjust samplesPerChunk from 120 to 220 Signed-off-by: Justin Lei <justin.lei@grafana.com>	2023-04-06 09:19:32 -07:00
Ganesh Vernekar	e709b0b36e	Merge pull request #12127 from codesome/ooo-mmap-replay Update OOO min/max time properly after replaying m-map chunks	2023-04-04 12:05:57 +05:30
Ganesh Vernekar	5588cab8b2	Merge pull request #12173 from bboreham/builder-no-empty-labels labels: simplify call to get Labels from Builder	2023-04-04 12:02:55 +05:30
Alex Le	1936868e9d	Allow populate block logic in compact to be overriden outside Prometheus (#11711 ) Signed-off-by: Alex Le <leqiyue@amazon.com> Signed-off-by: Alex Le <emoc1989@gmail.com>	2023-04-04 12:01:49 +05:30
Ganesh Vernekar	f55ab22179	Merge pull request #12186 from codesome/remove-file Remove mistakenly added file	2023-03-30 19:24:04 +05:30
Oleg Zaytsev	3ded84e649	Fix TestCancelCompactions on windows It seems that readOnlyDB was still opened which blocked the temp dir cleanup. Also changed the copy dir to be another TempDir instead of manually creating one. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2023-03-30 13:38:43 +02:00
Björn Rabenstein	ae42dd4c4a	Merge pull request #12179 from colega/fix-block-compaction-failed-when-shutting-down Fix block compaction failed when shutting down	2023-03-30 13:27:49 +02:00
Oleg Zaytsev	6e2905a4d4	Use zeropool.Pool to workaround SA6002 (#12189 ) * Use zeropool.Pool to workaround SA6002 I built a tiny library called https://github.com/colega/zeropool to workaround the SA6002 staticheck issue. While searching for the references of that SA6002 staticheck issues on Github first results was Prometheus itself, with quite a lot of ignores of it. This changes the usages of `sync.Pool` to `zeropool.Pool[T]` where a pointer is not available. Also added a benchmark for HeadAppender Append/Commit when series already exist, which is one of the most usual cases IMO, as I didn't find any. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> * Improve BenchmarkHeadAppender with more cases Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> * A little copying is better than a little dependency https://www.youtube.com/watch?v=PAAkCSZUG1c&t=9m28s Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> * Fix imports order Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> * Add license header Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> * Copyright should be on one of the first 3 lines Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> * Use require.Equal for testing I don't depend on testify in my lib, but here we have it available. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> * Avoid flaky test Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> * Also use zeropool for pointsPool in engine.go Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> --------- Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2023-03-29 20:34:34 +01:00
Ganesh Vernekar	b33a382646	Remove mistakenly added file It got added in https://github.com/prometheus/prometheus/pull/11992 Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-03-27 20:44:11 +05:30
Alan Protasio	6ddadd98b4	Optimization on `mergedStringIter` (#12132 ) Optimization on NewMergedStringIter Signed-off-by: Alan Protasio <alanprot@gmail.com>	2023-03-27 17:10:45 +05:30
Oleg Zaytsev	344c630857	Fix context.Canceled wrapping in compaction We need to make sure that `tsdb_errors.NewMulti` handles the errors.Is() calls properly, like it's done in grafana/dskit. Also we need to check that `errors.Is(err, context.Canceled)`, not that `err == context.Canceled`. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2023-03-23 11:10:00 +01:00
Oleg Zaytsev	2f32a9e3c3	Test compaction not failed during shutdown Test that blocks are not marked as "compaction failed" during shutdown. This shouldn't happen but this test currently fails. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2023-03-23 11:08:56 +01:00
Bryan Boreham	b987afa7ef	labels: simplify call to get Labels from Builder It took a `Labels` where the memory could be re-used, but in practice this hardly ever benefitted. Especially after converting `relabel.Process` to `relabel.ProcessBuilder`. Comparing the parameter to `nil` was a bug; `EmptyLabels` is not `nil` so the slice was reallocated multiple times by `append`. Lastly `Builder.Labels()` now estimates that the final size will depend on labels added and deleted. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2023-03-22 17:05:20 +00:00
Bryan Boreham	90b2f7a540	Merge pull request #12161 from codesome/update-comment tsdb: Fix a comment in tsdb/head_read.go	2023-03-22 17:02:28 +00:00
Vernon Miller	ca0abf26c5	Adds an affirmative log message for successful WAL repair (#12135 ) * Adds an affirmative log message for successful WAL repair Signed-off-by: Vernon Miller <vernon.miller@grafana.com> Signed-off-by: Vernon Miller <96601789+aldernero@users.noreply.github.com> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-03-21 19:33:43 +05:30
Ganesh Vernekar	1b7d973f14	tsdb: Fix a comment in tsdb/head_read.go Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-03-21 15:15:36 +05:30
Abhijit Mukherjee	8f6d5dcd45	Fix: getting rid of EncOOOXOR chunk encoding (#12111 ) Signed-off-by: mabhi <abhijit.mukherjee@infracloud.io>	2023-03-16 15:53:47 +05:30
Ganesh Vernekar	58a8d526e8	Merge pull request #11992 from codesome/no-reencode-chunk Do not re-encode head chunk for ChunkQuerier	2023-03-15 18:30:38 +05:30
Ganesh Vernekar	0a3f203c63	Update tests to not assume the chunk implementation Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-03-15 17:58:37 +05:30
Ganesh Vernekar	45b025898f	Add BenchmarkHeadChunkQuerier and BenchmarkHeadQuerier Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-03-15 17:58:31 +05:30
Ganesh Vernekar	0c0c2af7f5	Do not re-encode head chunk in ChunkQuerier Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-03-15 17:58:01 +05:30
Ganesh Vernekar	2af44f9558	tsdb: Update OOO min/max time properly after replaying m-map chunks Without this fix, if snapshots were enabled, and wbl goes missing between restarts, then TSDB does not recognize that there are ooo mmap chunks on disk and we cannot query them until those chunks are compacted into blocks. Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-03-13 13:14:00 +05:30
Ganesh Vernekar	1c3f1216b3	tsdb: Test querying after missing wbl with snapshots enabled If the snapshot was enabled with some ooo mmap chunks on disk, and wbl was removed between restarts, then we should still be able to query the ooo mmap chunks after a restart. This test shows that we are not able to query those ooo mmap chunks after a restart under this situation. Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-03-13 13:14:00 +05:30
Ganesh Vernekar	c9d06f2826	tsdb: Replay m-map chunk only when required M-map chunks replayed on startup are discarded if there was no WAL and no snapshot loaded, because there is no series created in the Head that it can map to. So only load m-map chunks from disk if there is either a snapshot loaded or there is WAL on disk. Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-03-13 13:13:42 +05:30
Ganesh Vernekar	6c008ec56a	Merge pull request #11962 from jesusvazquez/jvp/protect-new-compaction-head-from-uninitialized-wbl TSDB: Protect NewOOOCompactionHead from an uninitialized wbl	2023-03-13 10:52:03 +05:30
Đurica Yuri Nikolić	c9b85afd93	Making the number of CPUs used for WAL replay configurable (#12066 ) Adds `WALReplayConcurrency` as an option on tsdb `Options` and `HeadOptions`. If it is not set or set <=0, then `GOMAXPROCS` is used, which matches the previous behaviour. Signed-off-by: Yuri Nikolic <durica.nikolic@grafana.com>	2023-03-07 16:41:33 +00:00
ansalamdaniel	c1c444504e	Feat: metrics for head_chunks & wal folders (#12013 ) Signed-off-by: ansalamdaniel <ansalam.daniel@infracloud.io>	2023-03-02 15:25:56 +05:30
Rens Groothuijsen	d33eb3ab17	Automatically remove incorrect snapshot with index that is ahead of WAL (#11859 ) Signed-off-by: Rens Groothuijsen <l.groothuijsen@alumni.maastrichtuniversity.nl> Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-03-01 17:51:02 +05:30
Bryan Boreham	f34b2cede3	Remove microbenchmarks These benchmarks are all testing things related to what Prometheus does, so perhaps have some historical interest, but we should not retain them in the main repo. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2023-02-22 16:36:45 +00:00
Ganesh Vernekar	66da1d51fd	Merge pull request #12003 from codesome/redundant-chunk-access Remove unnecessary chunk fetch in Head queries	2023-02-22 12:57:38 +05:30
Ganesh Vernekar	d504c950a2	Remove unnecessary chunk fetch in Head queries `safeChunk` is only obtained from the `headChunkReader.Chunk` call where the chunk is already fetched and stored with the `safeChunk`. So, when getting the iterator for the `safeChunk`, we don't need to get the chunk again. Also removed a couple of unnecessary fields from `safeChunk` as a part of this. Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-02-22 12:21:12 +05:30
Vishal N	96ba6831ae	Observe delta in seconds prometheus_tsdb_sample_ooo_delta Signed-off-by: Vishal Nadagouda <vishalmn1996@gmail.com>	2023-02-21 18:55:09 +05:30
Jesus Vazquez	5c3f058755	Add unit test and also protect truncateOOO Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com>	2023-02-10 15:18:17 +01:00
Jesus Vazquez	f269077855	Protect NewOOOCompactionHead from an unitialized wbl Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com>	2023-02-10 13:00:29 +01:00
Justin Lei	af1d9e01c7	Refactor tsdbutil for tests/native histograms (#11948 ) * Add float histograms to ChunkFromSamplesGeneric Signed-off-by: Justin Lei <justin.lei@grafana.com> * Add GenerateSamples functions to tsdbutil Signed-off-by: Justin Lei <justin.lei@grafana.com> PR responses Signed-off-by: Justin Lei <justin.lei@grafana.com> --------- Signed-off-by: Justin Lei <justin.lei@grafana.com>	2023-02-10 17:09:33 +05:30
George Krajcsovits	1f0cc09579	Export single ith test histogram generation functions (#11911 ) * Export single ith test histogram generation functions Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com> * Do not set counter reset hint for non-gauge histograms individually Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com> * Apply suggestions from code review Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com> Signed-off-by: George Krajcsovits <krajorama@users.noreply.github.com> --------- Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com> Signed-off-by: George Krajcsovits <krajorama@users.noreply.github.com> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-02-01 16:23:38 +05:30
Ganesh Vernekar	8e8b718365	Merge pull request #11858 from fayzal-g/fix-chunks-metrics tsdb: when reading WAL, correctly update chunksRemoved and chunks metrics	2023-01-30 19:45:12 +05:30
beorn7	1cfc8f65a3	histograms: Return actually useful counter reset hints This is a bit more conservative than we could be. As long as a chunk isn't the first in a block, we can be pretty sure that the previous chunk won't disappear. However, the incremental gain of returning NotCounterReset in these cases is probably very small and might not be worth the code complications. Wwith this, we now also pay attention to an explicitly set counter reset during ingestion. While the case doesn't show up in practice yet, there could be scenarios where the metric source knows there was a counter reset even if it might not be visible from the values in the histogram. It is also useful for testing. Signed-off-by: beorn7 <beorn@grafana.com>	2023-01-25 16:57:21 +01:00
beorn7	57c18420ab	histograms: General readability tweaks - Adjust doc comments to go1.19 style. - Break down some overly long lines. - Minor doc comment tweaks and fixes. - Some renaming. Some rationales for the last point: I have renamed “interjections” into “inserts”, mostly because it is shorter, and the word shows up a lot by now (and the concept is cryptic enough to not obfuscate it even more with abbreviations). I have also tried to find more descriptive naming for the “compare spans” functions. Signed-off-by: beorn7 <beorn@grafana.com>	2023-01-19 13:26:42 +01:00
fayzal-g	cfa4ea53cc	Correctly update chunksRemoved and chunks metrics Signed-off-by: fayzal-g <fayzal.ghantiwala@grafana.com>	2023-01-18 10:58:48 +00:00
Ganesh Vernekar	6e560fe19b	tsdb: Avoid unnecessary allocation from 11779 Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-01-17 16:53:49 +05:30
Mingjie Shao	78d3c4e823	tsdb: Fixed typo in Histogram Signed-off-by: Mingjie Shao <com.jerryshao@jerryshao.com>	2023-01-16 18:13:45 +08:00
Ganesh Vernekar	cb2be6e62f	Merge pull request #11779 from codesome/memseries-ooo tsdb: Only initialise out-of-order fields when required	2023-01-16 10:58:05 +05:30
Jesus Vazquez	136956cca4	Attempt to append ooo sample at the end first (#11615 ) This is an optimization on the existing append in OOOChunk. What we've been doing so far is find the place inside the out-of-order slice where the new sample should go in and then place it there and move any samples to the right if necessary. This is OK but requires a binary search every time the slice is bigger than 0. The optimization is opinionated and suggests that although out-of-order samples can be out-of-order amongst themselves they'll probably be in order thus we can probably optimistically append at the end and if not do the binary search. OOOChunks are capped to 30 samples by default so this is a small optimization but everything adds up, specially if you handle many active timeseries with out-of-order samples. Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> Signed-off-by: Jesus Vazquez <jesusvazquez@users.noreply.github.com> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-01-13 19:00:50 +05:30
Marc Tudurí	721f33dbb0	histograms: Add remote-write support for Float Histograms (#11817 ) * adapt code.go and write_handler.go to support float histograms * adapt watcher.go to support float histograms * wip adapt queue_manager.go to support float histograms * address comments for metrics in queue_manager.go * set test cases for queue manager * use same counts for histograms and float histograms * refactor createHistograms tests * fix float histograms ref in watcher_test.go * address PR comments Signed-off-by: Marc Tuduri <marctc@protonmail.com>	2023-01-13 16:39:20 +05:30
Sebastian Rabenhorst	c057318578	agent: native histogram support (#11842 ) Signed-off-by: Sebastian Rabenhorst <sebastian.rabenhorst@shopify.com>	2023-01-12 11:13:44 -05:00
Ganesh Vernekar	38fa151a7c	tsdb: Only initialise out-of-order fields when required Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-01-12 20:29:16 +05:30
beorn7	6dcd03dbf3	tsdb: Add integer gauge histogram support This follows what #11783 has done for float gauge histograms. Signed-off-by: beorn7 <beorn@grafana.com>	2023-01-11 13:28:43 +01:00
Ganesh Vernekar	57bcbf1888	Merge pull request #11783 from codesome/gauge-histogram tsdb: Add gauge histogram support	2023-01-10 19:06:08 +05:30
Ganesh Vernekar	3c2ea91a83	tsdb: Test gauge float histograms Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-01-10 18:35:37 +05:30
Ganesh Vernekar	609b12d719	tsdb: Support gauge float histogram with recoding of chunk Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-01-10 17:48:09 +05:30
Ganesh Vernekar	8ad0d2d5d7	tsdb: Find union of two sets of histogram spans Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-01-10 17:43:33 +05:30
Ganesh Vernekar	d7f5129042	tsdb: Add logic to determine appendable gauge float histograms This is to check if a gauge histogram can be appended to the given chunk. If not, it tells what changes to make to the chunk and the histogram if possible. Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-01-10 17:43:33 +05:30
Ganesh Vernekar	a87e7e9e33	tsdb: Add counter reset hint to histograms and support in WAL Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-01-10 17:41:53 +05:30
Oleg Zaytsev	de93a279a0	Shortcut postings for matchers when empty postings are selected (#11813 ) * Add more benchmark cases * Add shortcuts for empty postings Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2023-01-10 15:21:49 +05:30
Ganesh Vernekar	fd89d7892c	Merge pull request #11809 from bboreham/dont-sort-postings-values tsdb: sort values for Postings only when required	2023-01-10 15:02:21 +05:30
Ganesh Vernekar	c94a41c4b2	Merge pull request #11785 from Fish-pro/erroris Use errors.Is to check for a specific error	2023-01-10 14:56:14 +05:30
György Krajcsovits	97626c9583	Fix comment Comment was not updated when code changed from labels to builder in #11717 Signed-off-by: György Krajcsovits <gyorgy.krajcsovits@grafana.com>	2023-01-08 16:29:02 +01:00
Björn Rabenstein	c49a28bb97	Merge pull request #11782 from codesome/floatappendabletest tsdb: Improve TestFloatHistogramChunkAppendable and TestHistogramChunkAppendable	2023-01-05 17:15:10 +01:00
Bryan Boreham	e61348d9f3	tsdb/index: fast-track postings for label="" We need to special-case ""="" too, which is used in some tests to mean "everything". Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2023-01-05 14:05:54 +00:00
Bryan Boreham	cf92cd2688	tsdb: sort values for Postings only when required In the head and in v1 postings on disk, it makes no difference whether postings are sorted. Only for v2 does the code step through in order. So, move the sorting to where it is required, and thus skip it entirely in the head. Label values in on-disk blocks are already sorted, but `slices.Sort` is very fast on already-sorted data so we don't bother checking. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2023-01-05 14:05:54 +00:00
Ganesh Vernekar	7ed1ddb338	tsdb: Improve TestHistogramChunkAppendable and add new cases Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2023-01-05 14:44:24 +05:30
Ganesh Vernekar	fa0f04bbc6	Merge pull request #11805 from bboreham/fix-benchmark-intersect tsdb/index: fix BenchmarkIntersect to do work on each loop	2023-01-04 18:19:14 +05:30
Bryan Boreham	3da2c99ffd	tsdb/index: don't call ExpandPostings in a benchmark This allocates memory for all the returned values, which skews the result. We aren't trying to benchmark `ExpandPostings`, so just step through all the values without storing them to consume them. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2023-01-03 15:26:29 +00:00
Bryan Boreham	4931983ca9	tsdb/index: make BenchmarkIntersect do work on each loop Previously all the postings constructed were consumed on the first iteration, so subsequent iterations did no work. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2023-01-03 15:25:38 +00:00
Fish-pro	6ed71a229e	Use errors.Is to check for a specific error Signed-off-by: Fish-pro <zechun.chen@daocloud.io>	2022-12-29 23:23:07 +08:00
Ganesh Vernekar	b42802af9a	tsdb: Improve TestFloatHistogramChunkAppendable and add new cases Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-12-28 21:07:47 +05:30
Ganesh Vernekar	c155c0e312	tsdb: Test staleness handling of FloatHistogram Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-12-28 14:48:56 +05:30
Ganesh Vernekar	2820e327db	tsdb: Add staleness handling for FloatHistogram Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-12-28 14:48:39 +05:30
Ganesh Vernekar	e555469ba1	tsdb: Remove isHistogramSeries from memSeries Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-12-28 14:31:55 +05:30
Marc Tudurí	9474610baf	Support FloatHistogram in TSDB (#11522 ) Extends Appender.AppendHistogram function to accept the FloatHistogram. TSDB supports appending, querying, WAL replay, for this new type of histogram. Signed-off-by: Marc Tudurí <marctc@protonmail.com> Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-12-28 14:25:07 +05:30
Bryan Boreham	1848623c77	tsdb: re-use iterator when stepping through chunks Saves memory allocations, hence reduces garbage-collection overheads. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-12-22 17:01:47 +00:00
Bryan Boreham	ccea61c7bf	Merge pull request #11717 from bboreham/labels-abstraction Add and use abstractions over labels.Labels	2022-12-20 17:23:39 +00:00
Ganesh Vernekar	6fd89a6fd2	Add chunk encoding for float histogram (#11716 ) Signed-off-by: Marc Tudurí <marctc@protonmail.com> Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Co-authored-by: Marc Tudurí <marctc@protonmail.com>	2022-12-20 15:33:32 +05:30
Bryan Boreham	10b27dfb84	Simplify IndexReader.Series interface Instead of passing in a `ScratchBuilder` and `Labels`, just pass the builder and the caller can extract labels from it. In many cases the caller didn't use the Labels value anyway. Now in `Labels.ScratchBuilder` we need a slightly different API: one to assign what will be the result, instead of overwriting some other `Labels`. This is safer and easier to reason about. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-12-19 15:22:09 +00:00
Bryan Boreham	4b6a4d1425	Update package tsdb tests for new labels.Labels type Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-12-19 15:22:09 +00:00
Bryan Boreham	ce2cfad0cb	Update package tsdb/record for new labels.Labels type Implement decoding via labels.ScratchBuilder, which we retain and re-use to reduce memory allocations. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-12-19 15:22:09 +00:00
Bryan Boreham	543c318ec2	Update package tsdb for new labels.Labels type Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-12-19 15:22:09 +00:00
Bryan Boreham	f0ec81badd	Update package tsdb/test for new labels.Labels type Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-12-19 15:22:09 +00:00
Bryan Boreham	14ad2e780b	Update package tsdb/agent for new labels.Labels type Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-12-19 15:22:09 +00:00
Bryan Boreham	a5bdff414b	Update package tsdb/index tests for new labels.Labels type Note in one cases we needed an extra copy of labels in case they change. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-12-19 15:22:09 +00:00
Bryan Boreham	d3d96ec887	tsdb/index: use ScratchBuilder to create Labels This necessitates a change to the `tsdb.IndexReader` interface: `index.Reader` is used from multiple goroutines concurrently, so we can't have state in it. We do retain a `ScratchBuilder` in `blockBaseSeriesSet` which is iterator-like. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-12-19 15:22:09 +00:00
Bryan Boreham	927a14b0e9	Update package tsdb/index for new labels.Labels type Incomplete - needs further changes to `Decoder.Series()`. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-12-19 15:22:09 +00:00
Bryan Boreham	89bf6e1df9	tsdb: Tidy up some test code Use simpler utility function to create Labels objects, making fewer assumptions about the data structure. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-12-15 19:39:46 +00:00
Bryan Boreham	0853250695	Review feedback Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-12-15 18:32:45 +00:00
Bryan Boreham	463f5cafdd	storage: re-use iterators to save garbage Re-use previous memory if it is already of the correct type. In `NewListSeries` we hoist the conversion to an interface value out so it only allocates once. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-12-15 18:32:45 +00:00
Bryan Boreham	f0866c0774	tsdb: optimise block series iterators Re-use previous memory if it is already of the correct type. Also turn two levels of function closure into a single object that holds the required data. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-12-15 18:32:45 +00:00
Bryan Boreham	3c7de69059	storage: allow re-use of iterators Patterned after `Chunk.Iterator()`: pass the old iterator in so it can be re-used to avoid allocating a new object. (This commit does not do any re-use; it is just changing all the method signatures so re-use is possible in later commits.) Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-12-15 18:32:45 +00:00
Julien Pivotto	475cfe8a6b	Merge remote-tracking branch 'origin/release-2.40' Signed-off-by: Julien Pivotto <roidelapluie@o11y.eu>	2022-12-14 11:22:01 +01:00
Ganesh Vernekar	db99fc43e4	Merge pull request #11632 from bboreham/improve-bbss tsdb: improve blockBaseSeriesSet scan	2022-12-14 15:05:27 +05:30
Ganesh Vernekar	54739a1465	Merge pull request #11674 from bboreham/fix-tsdb-test-mem tsdb tests: allocate more reasonable sample slice	2022-12-14 15:01:04 +05:30
beorn7	5f366e9b62	histograms: Improve tests and fix exposed bugs This adds negative buckets and access of float histograms to TestHistogramChunkSameBuckets and TestHistogramChunkBucketChanges. It also exercises a specific pattern of reusing an iterator (one where no access has happened). This exposes two bugs (where entries for positive buckets where used where the corresponding entries for negative buckets should have been used). One was fixed in #11627 (not merged), which triggered the work in this commit. This commit fixes both issues, so #11627 can be closed. It also simplifies the code in the histogramIterator.Next method that aims to recycle existing slice capacity. Furthermore, this is on top of the release-2.40 branch because we should probably cut a bugfix release for this. Signed-off-by: beorn7 <beorn@grafana.com>	2022-12-12 00:08:23 +01:00
Julien Pivotto	0b302f8a39	Merge pull request #11662 from prometheus/release-2.40 Merge back release-2.40 branch again	2022-12-06 17:30:51 +01:00
Bryan Boreham	9853888f9b	tsdb tests: allocate more reasonable sample slice Typical parameters are one hour by 1 minute step, where the function would allocate a slice of 3.6 million samples instead of 60. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-12-05 17:15:02 +00:00
Ganesh Vernekar	72a48321da	Merge pull request #11633 from pstibrany/populate-error Enhance "cannot populate chunk" error message to include source block ID	2022-12-02 16:28:52 +05:30
Ganesh Vernekar	b8b0d45d69	Fix reset of a histogram chunk iterator Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-11-30 17:50:05 +05:30
Julien Pivotto	0372e259ba	Merge pull request #11634 from prometheus/release-2.40 Merge release-2.40 branch into main	2022-11-29 15:54:58 +01:00
Bryan Boreham	6bdecf377c	Switch from 'sanity' to more inclusive lanuage (#9376 ) * Switch from 'sanity' to more inclusive lanuage "Removing ableist language in code is important; it helps to create and maintain an environment that welcomes all developers of all backgrounds, while emphasizing that we as developers select the most articulate, precise, descriptive language we can rather than relying on metaphors. The phrase sanity check is ableist, and unnecessarily references mental health in our code bases. It denotes that people with mental illnesses are inferior, wrong, or incorrect, and the phrase sanity continues to be used by employers and other individuals to discriminate against these people." From https://gist.github.com/seanmhanson/fe370c2d8bd2b3228680e38899baf5cc Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-11-28 17:09:18 +00:00
Peter Štibraný	af838ccf83	Include source block in error message when loading chunk fails. Signed-off-by: Peter Štibraný <pstibrany@gmail.com>	2022-11-28 09:12:54 +01:00
Bryan Boreham	1226922ff5	tsdb: improve blockBaseSeriesSet scan Inverting the test for chunks deleted by tombstones makes all three rejections consistent, and also avoids the case where a chunk is excluded but still causes `trimFront` or `trimBack` to be set. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-11-26 15:23:02 +00:00
Bryan Boreham	0c05f95e92	tsdb: use smaller allocation in blockBaseSeriesSet This reduces garbage, hence goes faster, when a short time range is required compared to the amount of chunks in the block. For example recording rules and alerts often look only at the last few minutes. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-11-26 14:56:22 +00:00
Ganesh Vernekar	ad79fb9f25	Do not error on empty chunk during iteration in populateWithDelChunkSeriesIterator Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-11-23 17:32:28 +05:30
Ganesh Vernekar	d0e683e26d	Add TestCompactHeadWithDeletion to test compaction failure after deletion Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-11-23 17:31:18 +05:30
Ganesh Vernekar	42633bd05c	Merge pull request #11485 from t00350320/prometheus-office GetRefByhash() will query a label's ref with hash value rather than lset.Hash().	2022-11-16 15:09:49 +01:00
tanghengjian	982007ecab	GetRefByhash will query a label's ref with hash value rather than lset.Hash(). Signed-off-by: tanghengjian <1040104807@qq.com>	2022-11-16 14:13:59 +01:00
Oleg Zaytsev	8553a98267	Optimize postings offset table reading (#11535 ) * Add BenchmarkOpenBlock * Use specific types when reading offset table Instead of reading a generic-ish []string, we can read a generic type which would be specifically labels.Label. This avoid allocating a slice that escapes to the heap, making it both faster and more efficient in terms of memory management. * Update error message for unexpected number of keys * s/posting offset table/postings offset table/ * Remove useless lastKey assignment * Use two []bytes vars, simplify Applied PR feedback: removed generics, moved the label indices reading to that specific test as we're not using it in production anyway, we're just testing what we've just built. Also using two []bytes variables for name and value that use the backing buffer instead of using strings, this reduces allocations a lot as we only copy them when we store them (this is optimized by the compiler). * Fix the dumb bug Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> Co-authored-by: Marco Pracucci <marco@pracucci.com>	2022-11-14 17:48:16 +01:00
Julien Pivotto	739494d81b	Fix alignment of atomic int64 (#11547 ) * Fix atomix int64 placement * Test main for 386 Signed-off-by: Julien Pivotto <roidelapluie@o11y.eu>	2022-11-09 11:18:49 +01:00
Ganesh Vernekar	fa6e05903f	Merge pull request #11447 from prometheus/sparsehistogram Add Support for Native Histograms This PR merges all the coding work that has been done in sparsehistogram branch over the last 1 year into main branch. Design doc on native histograms: https://docs.google.com/document/d/1cLNv3aufPZb3fNfaJgdaRBZsInZKKIHo9E6HinJVbpM/edit Some sneak peak: https://www.youtube.com/watch?v=T2GvcYNth9U	2022-10-26 17:10:46 -04:00
Viacheslav Panasovets	3d2e18bad5	Fix time.Since() in defer. Wrap in anonymous function (#11489 ) Function arguments in defer evaluated during definition of defer, not during execution Signed-off-by: Slavik Panasovets <slavik@google.com> Signed-off-by: Slavik Panasovets <slavik@google.com>	2022-10-26 00:26:12 +02:00
Björn Rabenstein	503ffba49a	chunkenc: Slightly optimize xorWrite/xoRead (#11476 ) With these changes, the "happy path" when the leading and trailing number of bits don't need an update, fewer operations are needed. The change is probably very marginal (no change in the benchmark added here, but the benchmark also doesn't cover non-changing values), and an argument could me made that avoiding pointers also has its benefits. However, I think that reducing the number of return values improves readability. Which convinced me that I should at least propose this. Signed-off-by: beorn7 <beorn@grafana.com>	2022-10-20 15:08:01 +05:30
Ganesh Vernekar	8ee4dfd40c	Fix the build after conflict resolution Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-10-12 17:59:42 +05:30
Ganesh Vernekar	648be89822	Merge remote-tracking branch 'upstream/main' into fix-conflict Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-10-12 14:20:02 +05:30
Ganesh Vernekar	8e29110949	Add/Improve unit tests for compaction with histogram (#11342 ) Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-10-12 13:31:12 +05:30
Ganesh Vernekar	507bfa46fd	Fix HistogramChunk's AtFloatHistogram() Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-10-12 10:38:13 +05:30
Signed-off-by: Jesus Vazquez	3362bf6d79	Fix merge conflicts Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-10-11 22:53:37 +05:30
Jesus Vazquez	775d90d5f8	TSDB: Rename wal package to wlog (#11352 ) The wlog.WL type can now be used to create a Write Ahead Log or a Write Behind Log. Before the prefix for wbl metrics was 'prometheus_tsdb_out_of_order_wal_' and has been replaced with 'prometheus_tsdb_out_of_order_wbl_'. Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> Signed-off-by: Jesus Vazquez <jesusvazquez@users.noreply.github.com> Co-authored-by: Ganesh Vernekar <15064823+codesome@users.noreply.github.com>	2022-10-10 20:38:46 +05:30
Sonali Rajput	9165aedb49	Fixed broken link in tsdb README.md Signed-off-by: Sonali Rajput <sonalirajput1088@gmail.com>	2022-10-07 16:20:20 +00:00
Jesus Vazquez	e934d0f011	Merge 'main' into sparsehistogram Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com>	2022-10-05 22:14:49 +02:00
Ganesh Vernekar	d0a6488c74	Update metrics for histograms Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-10-03 13:48:59 +05:30
Bryan Boreham	9b31adc4e8	tsdb: fix up sort call with faster slices.Sort (#11380 ) This call was added by PR #11075 merged before #11318 which changed all similar calls to `sort.Sort` into a faster one. Signed-off-by: Bryan Boreham <bjboreham@gmail.com> Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-10-01 12:55:40 -04:00
Bryan Boreham	3330d85ba8	Replace sort.Strings and sort.Ints with faster slices.Sort (#11318 ) Use new experimental package `golang.org/x/exp/slices`. slices.Sort works on values that are directly comparable, like ints, so avoids the overhad of an interface call to `.Less()`. Left tests unchanged, because they don't need the speed and it may be a cross-check that slices.Sort gives the same answer. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-09-30 20:03:56 +05:30
Bryan Boreham	7f2374b703	tsdb: faster postings sort with generic slices.Sort (#11054 ) Use new experimental package `golang.org/x/exp/slices`. Some of the speedup comes from comparing SeriesRef (which is an int64) directly rather than through an interface `.Less()` call; some comes from exp/slices using "pattern-defeating quicksort(pdqsort)". Signed-off-by: Bryan Boreham <bjboreham@gmail.com> Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-09-30 20:01:32 +05:30
Ganesh Vernekar	83d738e263	Fix 'invalid magic number 0' bug (#11338 ) Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-09-28 21:43:58 +05:30
Ganesh Vernekar	f34aeefe6e	Allow overlapping blocks by default (#11331 ) Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-09-28 19:17:54 +05:30
Robert Fratto	448cfda6c1	tsdb/agent: fix validation of default options (#9876 ) * tsdb/agent: fix application of defaults MaxTS was being incorrectly constrained to the truncation interval * add more tests to check validation * force MaxWALTime = MinWALTime if min > max Signed-off-by: Robert Fratto <robertfratto@gmail.com>	2022-09-27 19:41:43 +05:30
Bryan Boreham	d166da7b59	tsdb: stop saving a copy of last 4 samples in memSeries (#11296 ) * TSDB chunks: remove race between writing and reading Because the data is stored as a bit-stream, the last byte in the stream could change if the stream is appended to after an Iterator is obtained. Copy the last byte when the Iterator is created, so we don't have to read it later. Clarify in comments that concurrent Iterator and Appender are allowed, but the chunk must not be modified while an Iterator is created. (This was already the case, in order to copy the bstream slice header.) * TSDB: stop saving last 4 samples in memSeries This extra copy of the last 4 samples was introduced to avoid a race condition between reading the last byte of the chunk and writing to it. But now we have fixed that by having `bstreamReader` copy the last byte, we don't need to copy the last 4 samples. This change saves 56 bytes per series, which is very worthwhile when you have millions or tens of millions of series. * TSDB: tidy up stopIterator re-use Previous changes have left this code duplicating some lines; pull them out to a separate function and tidy up. * TSDB head_test: stop checking when iterators are wrapped The behaviour has changed so chunk iterators are only wrapped when transaction isolation requires them to stop short of the end. This makes tests fail which are checking the type. Tests should check the observable behaviour, not the type. Signed-off-by: Bryan Boreham <bjboreham@gmail.com> Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-09-27 19:32:05 +05:30
Bryan Boreham	ff00dee262	tsdb: turn off transaction isolation for head compaction (#11317 ) * tsdb: add a basic test for read/write isolation * tsdb: store the min time with isolationAppender So that we can see when appending has moved past a certain point in time. * tsdb: allow RangeHead to have isolation disabled This will be used when for head compaction. * tsdb: do head compaction with isolation disabled This saves a lot of work tracking appends done while compaction is ongoing. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-09-27 19:31:23 +05:30
Bryan Boreham	d0607435a2	tsdb: remove chunkRange and oooCapMax from memSeries (#11288 ) * tsdb: remove chunkRange from memSeries chunkRange is the (oddly-named) configured duration for the head block. We don't need a copy of this value per series. Pass it down where required, and remove the copy. The value in `Head` is only updated in `resetInMemoryState()`, which also discards all `memSeries`. * tsdb: remove oooCapMax from memSeries oooCapMax is the configured maximum capacity for an out-of-order chunk. Storing it per-series uses extra memory, and has surprising behaviour if users change the value in config - series created before the change will keep their old value. Instead, pass it down where required, and remove the per-series value. Signed-off-by: Bryan Boreham <bjboreham@gmail.com> Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-09-27 13:52:22 +05:30
Ganesh Vernekar	758e29258b	Add/Improve unit tests for compaction with histogram Part 2 (#11343 ) Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-09-23 14:01:10 +05:30
Jesus Vazquez	c1b669bf9b	Add out-of-order sample support to the TSDB (#11075 ) * Introduce out-of-order TSDB support This implementation is based on this design doc: https://docs.google.com/document/d/1Kppm7qL9C-BJB1j6yb6-9ObG3AbdZnFUBYPNNWwDBYM/edit?usp=sharing This commit adds support to accept out-of-order ("OOO") sample into the TSDB up to a configurable time allowance. If OOO is enabled, overlapping querying are automatically enabled. Most of the additions have been borrowed from https://github.com/grafana/mimir-prometheus/ Here is the list ist of the original commits cherry picked from mimir-prometheus into this branch: - `4b2198d7ec` - `2836e5513f` - `00b379c3a5` - `ff0dc75758` - `a632c73352` - `c6f3d4ab33` - `5e8406a1d4` - `abde1e0ba1` - `e70e769889` - `df59320886` Co-authored-by: Jesus Vazquez <jesus.vazquez@grafana.com> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com> Co-authored-by: Dieter Plaetinck <dieter@grafana.com> Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> * gofumpt files Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> * Add license header to missing files Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> * Fix OOO tests due to existing chunk disk mapper implementation Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> * Fix truncate int overflow Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> * Add Sync method to the WAL and update tests Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> * remove useless sync Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> * Update minOOOTime after truncating Head * Update minOOOTime after truncating Head Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix lint Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Add a unit test Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> * Load OutOfOrderTimeWindow only once per appender Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> * Fix OOO Head LabelValues and PostingsForMatchers Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> * Fix replay of OOO mmap chunks Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Remove unnecessary err check Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> * Prevent panic with ApplyConfig Signed-off-by: Ganesh Vernekar 15064823+codesome@users.noreply.github.com Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> * Run OOO compaction after restart if there is OOO data from WBL Signed-off-by: Ganesh Vernekar 15064823+codesome@users.noreply.github.com Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> * Apply Bartek's suggestions Co-authored-by: Bartlomiej Plotka <bwplotka@gmail.com> Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> * Refactor OOO compaction Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Address comments and TODOs - Added a comment explaining why we need the allow overlapping compaction toggle - Clarified TSDBConfig OutOfOrderTimeWindow doc - Added an owner to all the TODOs in the code Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> * Run go format Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> * Fix remaining review comments Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix tests Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Change wbl reference when truncating ooo in TestHeadMinOOOTimeUpdate Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> * Fix TestWBLAndMmapReplay test failure on windows Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Address most of the feedback Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Refactor the block meta for out of order Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix windows error Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix review comments Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com> Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Signed-off-by: Ganesh Vernekar 15064823+codesome@users.noreply.github.com Co-authored-by: Ganesh Vernekar <15064823+codesome@users.noreply.github.com> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com> Co-authored-by: Dieter Plaetinck <dieter@grafana.com> Co-authored-by: Oleg Zaytsev <mail@olegzaytsev.com> Co-authored-by: Bartlomiej Plotka <bwplotka@gmail.com>	2022-09-20 22:35:50 +05:30
Bryan Boreham	af6167df58	WAL loading: don't send empty buffers over chan (#11319 ) If some shards did not get any samples mapped, the buffer will be empty so sending it over the chan to `processWALSamples()` is a waste of time. This is especially likely now we are checking `minValidTime` before sending. Signed-off-by: Bryan Boreham <bjboreham@gmail.com> Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-09-20 19:43:30 +05:30
Ganesh Vernekar	2474c6fb2c	Error on amending histograms on append (#11308 ) * Error on amending histograms on append Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Rename Matches to Equals Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-09-19 13:10:30 +05:30
Bryan Boreham	d2701be53a	tsdb: remove chunk pool from memSeries (#11280 ) The chunk pool belongs to the head not to the series. Pass it down where required, and remove the copy of the pointer that `memSeries` was holding. `safeChunk` also needs to hold it, because in scenarios where it is used we don't have a reference to the head. However it was already holding `chunkDiskMapper` for the same reason, so no big change. Signed-off-by: Bryan Boreham <bjboreham@gmail.com> Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-09-15 13:22:09 +05:30
Björn Rabenstein	7ad36505d5	tsdb: Update comment about a possible space optimization (#11303 ) See also #11195 for the detailed reasoning. Signed-off-by: beorn7 <beorn@grafana.com> Signed-off-by: beorn7 <beorn@grafana.com>	2022-09-15 13:11:57 +05:30
Bryan Boreham	e49d596fb1	WAL loading: check sample time is valid earlier (#11307 ) There's no point splitting a sample onto the right shard and checking if the series needs to be re-mapped, if we're only going to discard it once it arrives at `ProcessWALSamples()`. Simply discard it earlier. Signed-off-by: Bryan Boreham <bjboreham@gmail.com> Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-09-15 12:36:57 +05:30
Ganesh Vernekar	d354f20c2a	Add a feature flag to control native histogram ingestion (#11253 ) * Add runtime config to control native histogram ingestion Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Make the config into a CLI flag Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-09-14 17:38:34 +05:30
Ganesh Vernekar	83e11014dd	Remove unnecessary tsdb/tsdbutil/buffer.go (#11302 ) Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-09-13 19:36:32 +05:30
Ganesh Vernekar	b2d01cbc57	Remove unnecessary code in encoding/decoding histograms (#11252 ) * Remove unnecessary code in encoding/decoding histograms Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix review comments Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-09-13 19:30:54 +05:30
Bryan Boreham	136f8b0ebb	tsdb: comment reason for isolation tracking reads (#11301 ) I find it useful to know why a restriction exists, to check whether that reason still applies, or in which other places it might apply. This is based on the note here: https://github.com/prometheus/prometheus/pull/9270#pullrequestreview-743820956 on the PR where the original comment was added. Signed-off-by: Bryan Boreham <bjboreham@gmail.com> Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-09-13 11:11:03 +02:00
Bryan Boreham	176fa38e76	tsdb: in tests use labels.FromStrings Replacing code which assumes the internal structure of `Labels`. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-09-09 13:34:49 +02:00
Bryan Boreham	0437dd7cee	tsdb/wal: in tests use labels.FromStrings Replacing code which assumes the internal structure of `Labels`. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-09-09 13:34:49 +02:00
Julien Pivotto	ec6c1f17d1	Update dependencies (#11287 ) Updating dependencies following CI changes and move to go 1.19 Signed-off-by: Julien Pivotto <roidelapluie@o11y.eu>	2022-09-09 13:28:55 +02:00
Julien Pivotto	96d5a32659	Update go to 1.19, set min version to 1.18 (#11279 ) * Update go to 1.19, set min version to 1.18 Signed-off-by: Julien Pivotto <roidelapluie@o11y.eu> * Update golangci-lint Signed-off-by: Julien Pivotto <roidelapluie@o11y.eu> Signed-off-by: Julien Pivotto <roidelapluie@o11y.eu>	2022-09-07 11:30:48 +02:00
Ganesh Vernekar	8f755f8f35	Extend createHead in tests to support histograms Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-08-29 20:18:02 +05:30
Ganesh Vernekar	f540c1dbd3	Add support for histograms in WAL checkpointing (#11210 ) * Add support for histograms in WAL checkpointing Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix review comments Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix tests Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-08-29 17:38:36 +05:30
Ganesh Vernekar	6383994f3e	Improve WAL/mmap chunks test for histograms (#11208 ) Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-08-29 16:21:32 +05:30
Ganesh Vernekar	d209a29a5b	Add unit test for histogram append and various querying scenarios (#11194 ) * Add unit test for histogram append and various querying scenarios Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * make lint happy Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix tests Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix review comments Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-08-29 15:35:03 +05:30
Abirdcfly	314aa45c2c	chore: remove duplicate word in comments (#11225 ) Signed-off-by: Abirdcfly <fp544037857@gmail.com> Signed-off-by: Abirdcfly <fp544037857@gmail.com>	2022-08-27 22:21:41 +02:00
Ganesh Vernekar	0f4e5196c4	Implement vertical compaction for native histograms (#11184 ) * Implement vertical compaction for native histograms Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix typo Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-08-22 19:04:39 +05:30
Xiaochao Dong	09187fb0cc	Replay WAL concurrently without blocking (#10973 ) * Replay WAL concurrently without blocking Signed-off-by: Xiaochao Dong (@damnever) <the.xcdong@gmail.com> * Resolve review comments Signed-off-by: Xiaochao Dong (@damnever) <the.xcdong@gmail.com> Signed-off-by: Xiaochao Dong (@damnever) <the.xcdong@gmail.com>	2022-08-17 19:23:57 +05:30
Łukasz Mierzwa	3196c98bc2	Reduce memSeries memory usage by decoupling metadata (#11152 ) Metadata was added recently but doesn't seem to be used much, at least as far as I could identify. Yet it's part of memSeries struct and so even when empty takes 48 bytes, which is a lot given that without it memSeries requires 224 bytes. This change turns it into a pointer on the struct, that get set only when metadata is actually set of given series. Signed-off-by: Łukasz Mierzwa <l.mierzwa@gmail.com> Signed-off-by: Łukasz Mierzwa <l.mierzwa@gmail.com>	2022-08-17 15:32:28 +05:30
beorn7	c9fd3c235d	Merge branch 'main' into sparsehistogram	2022-08-10 17:54:37 +02:00
Xiaochao Dong	1078081aec	Fix race condition when updating lastSeriesID during loading chunk snapshot (#11099 ) Signed-off-by: Xiaochao Dong (@damnever) <the.xcdong@gmail.com>	2022-08-04 13:39:14 +05:30
Levi Harrison	77a7af4461	Add histogram validation (#11052 ) * Add histogram validation Signed-off-by: Levi Harrison <git@leviharrison.dev> * Correct negative offset validation Signed-off-by: Levi Harrison <git@leviharrison.dev> * Address review comments Signed-off-by: Levi Harrison <git@leviharrison.dev> * Validation benchmark Signed-off-by: Levi Harrison <git@leviharrison.dev> * Add more checks Signed-off-by: Levi Harrison <git@leviharrison.dev> * Attempt to fix tests Signed-off-by: Levi Harrison <git@leviharrison.dev> * Fix stuff Signed-off-by: Levi Harrison <git@leviharrison.dev>	2022-07-29 09:52:49 -05:00
Levi Harrison	cb8582637a	Implement rollback for histograms (#11071 ) Signed-off-by: Levi Harrison <git@leviharrison.dev>	2022-07-29 14:18:53 +05:30
Bryan Boreham	00ec720c29	tsdb: extract functions to encode and decode labels (#11045 ) * tsdb/record: Extract functions to encode and decode labels Signed-off-by: Bryan Boreham <bjboreham@gmail.com> * tsdb: make use of Encode/Decode Labels Simplify the code by re-using routines from tsdb/record. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-07-26 20:12:00 +05:30
Paschalis Tsilias	a0f7c31c26	Fix type byte of WAL metadata records in docs (#11035 ) Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com>	2022-07-19 18:22:02 +05:30
Paschalis Tsilias	d1122e0743	Introduce TSDB changes for appending metadata to the WAL (#10972 ) * Append metadata to the WAL Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Remove extra whitespace; Reword some docstrings and comments Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Use RLock() for hasNewMetadata check Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Use single byte for metric type in RefMetadata Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Update proposed WAL format for single-byte type metadata Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Implementa MetadataAppender interface for the Agent Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Address first round of review comments Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Amend description of metadata in wal.md Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Correct key used to retrieve metadata from cache When we're setting metadata entries in the scrapeCace, we're using the p.Help(), p.Unit(), p.Type() helpers, which retrieve the series name and use it as the cache key. When checking for cache entries though, we used p.Series() as the key, which included the metric name _with_ its labels. That meant that we were never actually hitting the cache. We're fixing this by utiling the __name__ internal label for correctly getting the cache entries after they've been set by setHelp(), setType() or setUnit(). Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Put feature behind a feature flag Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Fix AppendMetadata docstring Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Reorder WAL format document Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Change error message of AppendMetadata; Fix access of s.meta in AppendMetadata Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Reuse temporary buffer in Metadata encoder Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Only keep latest metadata for each refID during checkpointing Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Fix test that's referencing decoding metadata Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Avoid creating metadata block if no new metadata are present Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Add tests for corrupt metadata block and relevant record type Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Fix CR comments Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Extract logic about changing metadata in an anonymous function Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Implement new proposed WAL format and amend relevant tests Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Use 'const' for metadata field names Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Apply metadata to head memSeries in Commit, not in AppendMetadata Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Add docstring and rename extracted helper in scrape.go Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Add tests for tsdb-related cases Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Fix linter issues vol1 Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Fix linter issues vol2 Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Fix Windows test by closing WAL reader files Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Use switch instead of two if statements in metadata decoding Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Fix review comments around TestMetadata* tests Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Add code for replaying WAL; test correctness of in-memory data after a replay Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Remove scrape-loop related code from PR Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Address first round of comments Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Simplify tests by sorting slices before comparison Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Fix test to use separate transactions Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Empty out buffer and record slices after encoding latest metadata Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Fix linting issue Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Update calculation for DroppedMetadata metric Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Rename MetadataAppender interface and AppendMetadata method to MetadataUpdater/UpdateMetadata Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Reuse buffer when encoding latest metadata for each series Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Fix review comments; Check all returned error values using two helpers Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Simplify use of helpers Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Satisfy linter Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com>	2022-07-19 10:58:52 +02:00
ZhangJian He	95b7d058ac	Fix markdown syntax in tsdb index.md (#11032 ) * Fix markdown syntax in tsdb index.md Signed-off-by: ZhangJian He <shoothzj@gmail.com> * Update tsdb/docs/format/index.md Co-authored-by: Ganesh Vernekar <15064823+codesome@users.noreply.github.com> Signed-off-by: ZhangJian He <shoothzj@gmail.com>	2022-07-18 04:28:52 -07:00
Levi Harrison	3d538351f6	Recognize exemplar record type in WAL watcher metrics (#11008 ) * Add exemplar record case Signed-off-by: Levi Harrison <git@leviharrison.dev> * recordType() -> Type.String() Signed-off-by: Levi Harrison <git@leviharrison.dev>	2022-07-18 15:54:11 +05:30
Levi Harrison	08f3ddb864	Sparse histogram remote-write support (#11001 )	2022-07-14 09:13:12 -04:00
beorn7	3ce988b031	Merge branch 'sparsehistogram' into beorn7/sparsehistogram	2022-07-13 18:07:54 +02:00
beorn7	28f028e938	Merge branch 'main' into sparsehistogram	2022-07-12 19:07:13 +02:00
beorn7	5d14046d28	tsdb: Fix chunk handling during appendHistogram Previously, the maxTime wasn't updated properly in case of a recoding happening. My apologies for reformatting many lines for line length. During the bug hunt, I tried to make things more readable in a reasonably wide editor window. Signed-off-by: beorn7 <beorn@grafana.com>	2022-07-06 18:44:53 +02:00
beorn7	642c5758ff	tsdb: Expose histogram append bug Signed-off-by: beorn7 <beorn@grafana.com>	2022-07-06 18:44:45 +02:00
beorn7	49be0784b4	tsdb: Fix chunk handling during histogram recoding Previously, the maxTime wasn't updated properly in case of a recoding happening. My apologies for reformatting many lines for line length. During the bug hunt, I tried to make things more readable in a reasonably wide editor window. Signed-off-by: beorn7 <beorn@grafana.com>	2022-07-06 14:34:02 +02:00
Jesus Vazquez	6cfe44d7fd	WaitUntilIdle optimize idling time (#10878 ) Relates to @bboreham optimization in https://github.com/prometheus/prometheus/pull/10859 Bryan did reduce the sleep time improving the deltas on the benchmark by quite a lot. However I've been working on a similar implementation for out of order and I noticed that we actually get into this method thousands of times. @ywwg had the brilliant idea of not always sleeping before the select but actually make it a case in the select so we only sleep if we need to. The benchmark deltas are amazing ``` ❯ benchstat old_implementation.txt new_implementation_using_time_after.txt name old time/op new time/op delta LoadWAL/batches=10,seriesPerBatch=100,samplesPerSeries=7200,exemplarsPerSeries=0,mmappedChunkT=0-8 521ms ±25% 253ms ± 6% -51.47% (p=0.008 n=5+5) LoadWAL/batches=10,seriesPerBatch=100,samplesPerSeries=7200,exemplarsPerSeries=36,mmappedChunkT=0-8 773ms ± 3% 369ms ±31% -52.23% (p=0.008 n=5+5) LoadWAL/batches=10,seriesPerBatch=100,samplesPerSeries=7200,exemplarsPerSeries=72,mmappedChunkT=0-8 592ms ±28% 297ms ±28% -49.80% (p=0.008 n=5+5) LoadWAL/batches=10,seriesPerBatch=100,samplesPerSeries=7200,exemplarsPerSeries=360,mmappedChunkT=0-8 547ms ± 2% 999ms ±187% ~ (p=0.690 n=5+5) LoadWAL/batches=10,seriesPerBatch=10000,samplesPerSeries=50,exemplarsPerSeries=0,mmappedChunkT=0-8 11.3s ± 4% 1.3s ±44% -88.48% (p=0.008 n=5+5) LoadWAL/batches=10,seriesPerBatch=10000,samplesPerSeries=50,exemplarsPerSeries=2,mmappedChunkT=0-8 11.1s ± 1% 1.2s ±20% -89.08% (p=0.008 n=5+5) LoadWAL/batches=10,seriesPerBatch=1000,samplesPerSeries=480,exemplarsPerSeries=0,mmappedChunkT=0-8 1.24s ± 3% 0.18s ± 7% -85.76% (p=0.008 n=5+5) LoadWAL/batches=10,seriesPerBatch=1000,samplesPerSeries=480,exemplarsPerSeries=2,mmappedChunkT=0-8 1.24s ± 2% 0.18s ± 5% -85.24% (p=0.008 n=5+5) LoadWAL/batches=10,seriesPerBatch=1000,samplesPerSeries=480,exemplarsPerSeries=5,mmappedChunkT=0-8 1.23s ± 5% 0.27s ±33% -77.73% (p=0.008 n=5+5) LoadWAL/batches=10,seriesPerBatch=1000,samplesPerSeries=480,exemplarsPerSeries=24,mmappedChunkT=0-8 1.28s ± 1% 0.36s ± 7% -71.51% (p=0.008 n=5+5) LoadWAL/batches=100,seriesPerBatch=1000,samplesPerSeries=480,exemplarsPerSeries=0,mmappedChunkT=3800-8 12.1s ± 1% 3.1s ± 6% -74.33% (p=0.008 n=5+5) LoadWAL/batches=100,seriesPerBatch=1000,samplesPerSeries=480,exemplarsPerSeries=2,mmappedChunkT=3800-8 12.1s ± 1% 3.4s ± 4% -71.94% (p=0.008 n=5+5) LoadWAL/batches=100,seriesPerBatch=1000,samplesPerSeries=480,exemplarsPerSeries=5,mmappedChunkT=3800-8 12.1s ± 1% 3.8s ±17% -68.35% (p=0.008 n=5+5) LoadWAL/batches=100,seriesPerBatch=1000,samplesPerSeries=480,exemplarsPerSeries=24,mmappedChunkT=3800-8 12.4s ± 1% 4.0s ±18% -67.71% (p=0.008 n=5+5) ``` Benchmarked on Linux ``` goos: linux goarch: amd64 pkg: github.com/prometheus/prometheus/tsdb cpu: 11th Gen Intel(R) Core(TM) i7-1165G7 @ 2.80GHz ``` Signed-off-by: Jesus Vazquez <jesus.vazquez@grafana.com>	2022-06-30 15:00:04 +02:00
Julien Pivotto	bacd776356	Merge pull request #10907 from damnever/fix/panic Fix panic if series is not found when deleting series	2022-06-30 11:23:08 +02:00
Peter Štibraný	ffc60d8397	Reduce chunk write queue memory usage 2 (#10874 ) * Job queue This PR reimplements chan chunkWriteJob with custom buffered queue that should use less memory, because it doesn't preallocate entire buffer for maximum queue size at once. Instead it allocates individual "segments" with smaller size. As elements are added to the queue, they fill individual segments. When elements are removed from the queue (and segments), empty segments can be thrown away. This doesn't change memory usage of the queue when it's full, but should decrease its memory footprint when it's empty (queue will keep max 1 segment in such case). Signed-off-by: Peter Štibraný <pstibrany@gmail.com> * Modify test to work with low resolution timer. Signed-off-by: Peter Štibraný <pstibrany@gmail.com> * Improve comments. Signed-off-by: Peter Štibraný <pstibrany@gmail.com>	2022-06-29 17:51:27 +05:30
Xiaochao Dong (@damnever)	6b042da2d8	Fix panic if series is not found when deleting series Signed-off-by: Xiaochao Dong (@damnever) <the.xcdong@gmail.com>	2022-06-24 15:55:32 +08:00
Steve Azzopardi	04fe2c9522	fix(tsdb): inc mmap corruption counter on mmap out of sequence error (#10406 ) What --- When we see out of sequence chunks increase the chunk corruption counter to indicate that one of the chunks was corrupted. Reference: https://github.com/prometheus/prometheus/pull/10406#issuecomment-1142595527 Signed-off-by: Steve Azzopardi <steveazz@outlook.com>	2022-06-22 14:03:12 +05:30
Peter Štibraný	03a2313f7a	Reduce chunk write queue memory usage (#10873 ) * dont waste space on the chunkRefMap * add time factor * add comments * better readability * add instrumentation and more comments * formatting * uppercase comments * Address review feedback. Renamed "free" to "shrink" everywhere, updated comments and threshold to 1000. * double space Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> Co-authored-by: Peter Štibraný <pstibrany@gmail.com> Co-authored-by: Mauro Stettler <mauro.stettler@gmail.com>	2022-06-17 13:11:39 +05:30
Bryan Boreham	9f77d23889	tsdb: commit data periodically in CreateBlock (#10788 ) To avoid building up data in memory, commit and make a new appender periodically. The number `commitAfter = 10000` was chosen arbitrarily; testing with 10x more or less gives slightly worse results. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-06-17 11:26:19 +05:30
Łukasz Mierzwa	d65f037def	Don't increment prometheus_tsdb_compactions_failed_total when context is canceled (#10772 ) When restarting Prometheus I sometimes see: caller=db.go:832 level=error component=tsdb msg="compaction failed" err="compact head: persist head block: 2 errors: populate block: context canceled; context canceled" And prometheus_tsdb_compactions_failed_total metric gets incremented. This makes it more difficult to write alerts based on prometheus_tsdb_compactions_failed_total metric since any restart can trigger it. Signed-off-by: Łukasz Mierzwa <l.mierzwa@gmail.com>	2022-06-17 11:21:43 +05:30
beorn7	095b6c93dd	Merge branch 'main' into sparsehistogram	2022-06-14 14:27:35 +02:00
Bryan Boreham	542b9ecdbd	tsdb: reduce sleep time when reading WAL (#10859 ) The code sleeps for a short time to allow goroutines to finish, however it seems the duration can be reduced a lot, speeding up the reading process. I checked using some WAL data from production, and the queue is almost always empty at the time we enter `waitForIdle()` so there is no danger of spinning in the tight loop. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-06-12 11:54:11 +05:30
songjiayang	c2af0de522	make sure response error when TOC parse failed Signed-off-by: songjiayang <songjiayang1@gmail.com>	2022-06-12 08:06:14 +08:00
beorn7	40ad5e284a	Merge branch 'main' into beorn7/sparsehistogram	2022-06-09 20:50:30 +02:00
Bryan Boreham	9f79a6f4b5	tsdb: faster CRC check by avoiding allocations (#10789 ) Instead of creating a new hashing object every time, call `crc32.Checksum` which computes the answer without allocations. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-06-08 08:00:59 +05:30
Matej Gera	1dd247f68b	Remote Write: Rename confusing `walDir` parameter to `dir` (#10464 ) * Rename walDir parameter to dir Signed-off-by: Matej Gera <matejgera@gmail.com> * Improve NewQueueManager comment Signed-off-by: Matej Gera <matejgera@gmail.com>	2022-05-30 21:45:30 -07:00
David Leadbeater	57f4aab27d	Update godoc links and remove note about TSDB versioning (#10754 ) Signed-off-by: David Leadbeater <dgl@dgl.cx>	2022-05-26 18:34:43 +10:00
maizige	10b677b826	fix typo (#10696 ) Update doc comment Signed-off-by: gemaizi <864321211@qq.com>	2022-05-25 18:01:45 +02:00
Filip Petkovski	d3cb39044e	Fix typo in symbol table size exceeded error message (#10746 ) This commit fixes a typo when reporting an error that the the symbols table size has been exceeded. Signed-off-by: Filip Petkovski <filip.petkovsky@gmail.com>	2022-05-25 10:40:36 +02:00
Julien Pivotto	6e3a0efe40	Make necessary change to compile promql parser to wasm (#10683 ) Signed-off-by: Julien Pivotto <roidelapluie@o11y.eu>	2022-05-12 09:12:05 +02:00
Matthias Rampke	78f2645787	test(tsdb): break up repeated test to avoid timeout (#10671 ) On macOS, the TestTombstoneCleanRetentionLimitsRace performs very poorly. It takes more than a second to write out one block, and as it writes 400 of them, we run into the 10-minute test timeout frequently. While this doesn't fix the actual performance issue, breaking each iteration into a subtest makes the test pass reliably (because each iteration comfortably finishes in under a minute). Related report: https://groups.google.com/g/prometheus-developers/c/jxQ6Ayg6VJ4/m/03H_DS9PDAAJ Signed-off-by: Matthias Rampke <matthias@prometheus.io>	2022-05-09 00:39:26 +02:00
Łukasz Mierzwa	88f9b248b4	Correctly format error message (#10669 ) Signed-off-by: Łukasz Mierzwa <l.mierzwa@gmail.com>	2022-05-06 00:42:31 +02:00
Bryan Boreham	4b9f248e85	unit tests: make all Labels sorted alphabetically (#10532 ) "Labels is a sorted set of labels. Order has to be guaranteed upon instantiation." says the comment, so fix all the tests that break this rule. For `BenchmarkLabelValuesWithMatchers()` and `BenchmarkHeadLabelValuesWithMatchers()` the amount of work done changes significantly if you put the labels in order, because all series refs get neatly partitioned by the `tens` label, so I renamed the labels to maintain the previous behaviour. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-05-04 23:41:36 +02:00
beorn7	654c07783c	Fix deprecation Signed-off-by: beorn7 <beorn@grafana.com>	2022-05-04 13:43:23 +02:00
beorn7	3bc711e333	Merge branch 'main' into sparsehistogram	2022-05-04 13:37:13 +02:00
Matthieu MOREL	e2ede285a2	refactor: move from io/ioutil to io and os packages (#10528 ) * refactor: move from io/ioutil to io and os packages * use fs.DirEntry instead of os.FileInfo after os.ReadDir Signed-off-by: MOREL Matthieu <matthieu.morel@cnp.fr>	2022-04-27 11:24:36 +02:00
Oleg Zaytsev	af0f6da5cb	Fix chunk overflow appending samples at a variable rate (#10607 ) * Add a test with variable samples rate append This test overflows the chunk created in memseries, and the total amount of samples in the (only) mmapped chunk is 29, instead of the 65565 appended ones. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> * Cut new chunk when rate prediction was wrong When appending samples at a slow rate, and then appending at a higher rate, the prediction we made to cut a new chunk is no longer valid. Sometimes this can even cause an overflow in the chunk, if more samples than uint16 can hold are appended. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> * Improve comment on 2samplesPerChunk Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> Assert that all chunks have less than 240 samples Also, trigger new chunk at 240, not at more than 240 Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2022-04-20 14:54:20 +02:00
Paschalis Tsilias	40c1efe8bc	tsdb/agent: Ignore duplicate exemplars (#10595 ) * tsdb/agent: Ignore duplicate exemplars Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Make each exemplar unique in TestCommit Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Re-Trigger CI for Windows and UI-related steps Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Change test comment to properly re-trigger pipeline Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Defer Close() calls for test agent and segment reader Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com>	2022-04-18 11:41:04 -04:00
Julien Pivotto	685ce9964d	Merge pull request #10599 from prometheus/release-2.35 Merge back release 2.35	2022-04-15 00:10:06 +02:00
Robert Fratto	286dfc70b7	tsdb/agent: port grafana/agent#676 (#10587 ) * tsdb/agent: port grafana/agent#676 grafana/agent#676 fixed an issue where a loading a WAL with multiple segments may result in ref ID collision. The starting ref ID for new series should be set to the highest ref ID across all series records from all WAL segments. This fixes an issue where the starting ref ID was incorrectly set to the highest ref ID found in the newest segment, which may not have any ref IDs at all if no series records have been appended to it yet. Signed-off-by: Robert Fratto <robertfratto@gmail.com> * tsdb/agent: update terminology (s/ref ID/nextRef) Signed-off-by: Robert Fratto <robertfratto@gmail.com>	2022-04-14 10:27:06 +02:00
chavacava	0b41fd6e71	Fix data races in WAL replay (#10571 ) Signed-off-by: chavacava <salvadorcavadini+github@gmail.com>	2022-04-12 16:00:20 +05:30
beorn7	7ee1836ef5	Merge branch 'main' into sparsehistogram	2022-04-05 18:31:19 +02:00
Bryan Boreham	2c1be4df7b	tsdb: more efficient sorting of postings read from WAL at startup (#10500 ) * tsdb: avoid slice-to-interface allocation in EnsureOrder This is pulling the `seriesRefSlice` out of the loop, so the compiler doesn't allocate a new one on the heap every time. Signed-off-by: Bryan Boreham <bjboreham@gmail.com> * tsdb: use pointer type in Pool for EnsureOrder As noted by staticcheck, Pool prefers the objects in the pool to have pointer type. This is a little more fiddly to code, but avoids allocation of a wrapper object every time a slice is put into the pool. Removed a comment that said fixing this has a performance penalty: not borne out by benchmarks. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-03-30 15:10:19 +05:30
Wilbert Guo	83a2e52bc2	Add SyncForState Implementation for Ruler HA (#10070 ) * continuously syncing activeAt for alerts Signed-off-by: Yijie Qin <qinyijie@amazon.com> Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> * add import Signed-off-by: Yijie Qin <qinyijie@amazon.com> Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> * Refactor SyncForState and add unit tests Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> * Format code Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> * Add hook for syncForState Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> Fix go lint Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> Refactor syncForState override implementation Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> Add syncForState override func as argument to Update() Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> Fix go formatting Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> Fix circleci test errors Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> Remove overrideFunc as argument to run() Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> * remove the syncForState Signed-off-by: Yijie Qin <qinyijie@amazon.com> * use the override function to decide if need to replace the activeAt or not Signed-off-by: Yijie Qin <qinyijie@amazon.com> * fix test case Signed-off-by: Yijie Qin <qinyijie@amazon.com> * fix format Signed-off-by: Yijie Qin <qinyijie@amazon.com> * Trigger build Signed-off-by: Yijie Qin <qinyijie@amazon.com> * fixing comments Signed-off-by: Yijie Qin <qinyijie@amazon.com> * return the result of map of alerts instead of single one Signed-off-by: Yijie Qin <qinyijie@amazon.com> * upper case the QueryforStateSeries Signed-off-by: Yijie Qin <qinyijie@amazon.com> * use a more generic rule group post process function type Signed-off-by: Yijie Qin <qinyijie@amazon.com> * fix indentation Signed-off-by: Yijie Qin <qinyijie@amazon.com> * fix gofmt Signed-off-by: Yijie Qin <qinyijie@amazon.com> * fix lint Signed-off-by: Yijie Qin <qinyijie@amazon.com> * fixing naming Signed-off-by: Yijie Qin <qinyijie@amazon.com> * fix comments Signed-off-by: Yijie Qin <qinyijie@amazon.com> * add the lastEvalTimestamp as parameter Signed-off-by: Yijie Qin <qinyijie@amazon.com> * fmt Signed-off-by: Yijie Qin <qinyijie@amazon.com> * change funcType to func Signed-off-by: Yijie Qin <qinyijie@amazon.com> Co-authored-by: Yijie Qin <qinyijie@amazon.com> Co-authored-by: Yijie Qin <63399121+qinxx108@users.noreply.github.com>	2022-03-29 02:16:46 +02:00
Howie	1291ec7185	deleting .tmp WAL files on startup (#10317 ) fix issue #10245 Signed-off-by: lihaowei <haoweili35@gmail.com> * minor changes Signed-off-by: lihaowei <haoweili35@gmail.com> * review changes Signed-off-by: lihaowei <haoweili35@gmail.com> * minor changes Signed-off-by: lihaowei <haoweili35@gmail.com>	2022-03-24 16:14:14 +05:30
beorn7	f9c411604d	Fix spelling errors Signed-off-by: beorn7 <beorn@grafana.com>	2022-03-22 16:02:13 +01:00
beorn7	4210aac74a	Merge branch 'main' into sparsehistogram	2022-03-22 14:47:42 +01:00
Chris Marchbanks	c1387494dd	Merge pull request #10452 from prometheus/release-2.34 Merge Release 2.34 into main	2022-03-15 12:32:18 -06:00
songjiayang	8d8be43824	Update wal.md (#10442 ) update exemplar record type Signed-off-by: songjiayang <songjiayang1@gmail.com>	2022-03-15 22:33:45 +05:30
Mauro Stettler	b025390cb4	Disable chunk write queue by default, allow user to configure the exact size (#10425 ) * Disable chunk write queue by default Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * update flag description Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com>	2022-03-11 17:26:59 +01:00
Łukasz Mierzwa	da23c4649a	Enable misspell check in golangci-lint (#10393 ) Signed-off-by: Łukasz Mierzwa <l.mierzwa@gmail.com>	2022-03-03 18:11:19 +01:00
Łukasz Mierzwa	a4317bf0ec	Run gofumpt on all files (#10392 ) * Run gofumpt on all files Getting golangci-lint errors when building on my laptop, possibly because I have newer version of gofumpt then what it was formatted with. Run gofumpt -w -extra on all files as it will be needed in the future anyway. * Update golangci-lint to v1.44.2 v1.44.0 upgraded gofumpt so bumping version in CI will help keep formatting correct for everyone * Address golangci-lint error Getting 'error-strings: error strings should not be capitalized or end with punctuation or a newline' from revive here. Drop new line. Signed-off-by: Łukasz Mierzwa <l.mierzwa@gmail.com>	2022-03-03 17:21:05 +01:00
cui fliter	c9b56d1a49	all: fix some typos (#10389 ) Signed-off-by: cuishuang <imcusg@gmail.com>	2022-03-03 12:03:07 +00:00
Ganesh Vernekar	4cc25c0cb0	Fix panic on query when m-map replay fails with snapshot enabled (#10348 ) * Fix panic on query when m-map replay fails with snapshot enabled Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix review Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix flake Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-02-25 08:53:40 -07:00
Björn Rabenstein	d1edb006c1	Merge pull request #10341 from prometheus/release-2.33 Merge release-2.33 forward into main	2022-02-22 22:51:05 +01:00
Ganesh Vernekar	24827782cb	Fix panics when m-mapping head chunks (#10316 ) * Fix panics when m-mapping head chunks Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix review comments Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix reviews Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-02-22 20:35:15 +05:30
Dieter Plaetinck	aa8874bc56	clarify Head.appendableMinValidTime (#10303 ) Signed-off-by: Dieter Plaetinck <dieter@grafana.com>	2022-02-17 16:30:48 +05:30
lwangrabbit	9fde6edbf5	tsdb/wal: Move comment of w.writer.Append(...) to the WriteTo interface (#10198 ) Signed-off-by: wanglipeng <wanglipeng@huayun.com> Co-authored-by: wanglipeng <wanglipeng@huayun.com>	2022-01-30 22:14:16 -08:00
Eng Zer Jun	3e67654d37	refactor: use `T.TempDir()` and `B.TempDir` to create temporary directory The directory created by `T.TempDir()` and `B.TempDir()` is automatically removed when the test and all its subtests complete. Reference: https://pkg.go.dev/testing#T.TempDir Reference: https://pkg.go.dev/testing#B.TempDir Signed-off-by: Eng Zer Jun <engzerjun@gmail.com>	2022-01-22 18:57:30 +08:00
Robert Fratto	b71a6dbbd1	tsdb/agent: Fix deadlock from simultaneous GC and write (#10166 ) * tsdb/agent: Fix deadlock from simultaneous GC and write This commit fixes a potential deadlock where storing in-memory series references could deadlock with a WAL GC cycle. Signed-off-by: Robert Fratto <robertfratto@gmail.com> * add missing license header Signed-off-by: Robert Fratto <robertfratto@gmail.com> * order local imports Signed-off-by: Robert Fratto <robertfratto@gmail.com> * align deadlock testing with discovery/manager_test.go method Also prevents GCs from running concurrently, which could also cause a deadlock (even though it's currently impossible for two GCs to run concurrently). Signed-off-by: Robert Fratto <robertfratto@gmail.com>	2022-01-19 20:23:06 +05:30
Mauro Stettler	bf959b36cb	Nits after PR 10051 merge (#10159 ) Signed-off-by: Marco Pracucci <marco@pracucci.com> Co-authored-by: Marco Pracucci <marco@pracucci.com>	2022-01-19 20:20:35 +05:30
Ganesh Vernekar	129ed4ec8b	Fix Example() function in TSDB (#10153 ) * Fix Example() function in TSDB Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix tests Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-01-11 17:24:03 +05:30
Mauro Stettler	0df3489275	Write chunks via queue, predicting the refs (#10051 ) * Write chunks via queue, predicting the refs Our load tests have shown that there is a latency spike in the remote write handler whenever the head chunks need to be written, because chunkDiskMapper.WriteChunk() blocks until the chunks are written to disk. This adds a queue to the chunk disk mapper which makes the WriteChunk() method non-blocking unless the queue is full. Reads can still be served from the queue. Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * address PR feeddback Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * initialize metrics without .Add(0) Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * change isRunningMtx to normal lock Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * do not re-initialize chunkrefmap Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * update metric outside of lock scope Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * add benchmark for adding job to chunk write queue Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * remove unnecessary "success" var Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * gofumpt -extra Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * avoid WithLabelValues call in addJob Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * format comments Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * addressing PR feedback Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * rename cutExpectRef to cutAndExpectRef Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * use head.Init() instead of .initTime() Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * address PR feedback Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * PR feedback Co-authored-by: Ganesh Vernekar <15064823+codesome@users.noreply.github.com> Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * update test according to PR feedback Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * replace callbackWg -> awaitCb Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * better test of truncation with empty files Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * replace callbackWg -> awaitCb Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> Co-authored-by: Ganesh Vernekar <15064823+codesome@users.noreply.github.com>	2022-01-10 13:36:45 +00:00
Oleg Zaytsev	a83d46ee9c	Tidy postingsWithIndexHeap (#10123 ) Unexported postingsWithIndexHeap's methods that don't need to be exported, and added detailed comments. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2022-01-06 16:03:44 +05:30
Bryan Boreham	82860a770c	tsdb: use simpler map key to improve exemplar ingest performance (#10111 ) * tsdb: fix exemplar benchmarks Go benchmarks are expected to do an amount of work that varies with the `b.N` parameter. Previously these benchmarks would report a result like 0.01 ns/op, which is nonsense. Signed-off-by: Bryan Boreham <bjboreham@gmail.com> * tsdb: use simpler map key to improve exemplar perf Prometheus holds an index of exemplars so it can discard the oldest one for a series when a new one is added. Since the keys are not for human eyes, we can use a simpler format and save the effort of quoting label values. Signed-off-by: Bryan Boreham <bjboreham@gmail.com> * Exemplars: allocate index map with estimated size This avoids Go having to re-size the map several times as it grows. 16 exemplars per series is a guess; if it is too low then the map will be sparse, while if it is too high then the map will have to resize once or twice. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-01-06 15:58:58 +05:30
Oleg Zaytsev	701545286d	Pop intersected postings heap without popping (#10092 ) See this comment for detailed explanation: https://github.com/prometheus/prometheus/pull/9907#issuecomment-1002189932 TL;DR: if we don't call Pop() on the heap implementation, we don't need to return our param as an `interface{}` so we save an allocation. This would be popped for every label value, so it can be thousands of saved allocations here (see benchmarks). Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2022-01-05 16:16:43 +05:30
Peter Štibraný	e51a17b501	CompactBlockMetas should produce correct mint/maxt for overlapping blocks. (#10108 ) Signed-off-by: Peter Štibraný <pstibrany@gmail.com>	2022-01-05 15:10:00 +05:30
Oleg Zaytsev	3947238ce0	Label values with matchers by intersecting postings (#9907 ) * LabelValues w/matchers by intersecting postings Instead of iterating all matched series to find the values, this checks if each one of the label values is present in the matched series (postings). Pending to be benchmarked. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> * Benchmark labelValuesWithMatchers name old time/op new time/op Querier/Head/labelValuesWithMatchers/i_with_n="1" 157ms ± 0% 48ms ± 0% Querier/Head/labelValuesWithMatchers/i_with_n="^.+$" 1.80s ± 0% 0.46s ± 0% Querier/Head/labelValuesWithMatchers/i_with_n="1",j!="foo" 144ms ± 0% 57ms ± 0% Querier/Head/labelValuesWithMatchers/i_with_n="1",i=~"^.$",j!="foo" 304ms ± 0% 111ms ± 0% Querier/Head/labelValuesWithMatchers/n_with_j!="foo" 761ms ± 0% 164ms ± 0% Querier/Head/labelValuesWithMatchers/n_with_i="1" 6.11µs ± 0% 6.62µs ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="1" 117ms ± 0% 62ms ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="^.+$" 1.44s ± 0% 0.24s ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="1",j!="foo" 92.1ms ± 0% 70.3ms ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="1",i=~"^.$",j!="foo" 196ms ± 0% 115ms ± 0% Querier/Block/labelValuesWithMatchers/n_with_j!="foo" 1.23s ± 0% 0.21s ± 0% Querier/Block/labelValuesWithMatchers/n_with_i="1" 1.06ms ± 0% 0.88ms ± 0% name old alloc/op new alloc/op Querier/Head/labelValuesWithMatchers/i_with_n="1" 29.5MB ± 0% 26.9MB ± 0% Querier/Head/labelValuesWithMatchers/i_with_n="^.+$" 46.8MB ± 0% 251.5MB ± 0% Querier/Head/labelValuesWithMatchers/i_with_n="1",j!="foo" 29.5MB ± 0% 22.3MB ± 0% Querier/Head/labelValuesWithMatchers/i_with_n="1",i=~"^.$",j!="foo" 46.8MB ± 0% 23.9MB ± 0% Querier/Head/labelValuesWithMatchers/n_with_j!="foo" 10.3kB ± 0% 138535.2kB ± 0% Querier/Head/labelValuesWithMatchers/n_with_i="1" 5.54kB ± 0% 7.09kB ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="1" 39.1MB ± 0% 28.5MB ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="^.+$" 287MB ± 0% 253MB ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="1",j!="foo" 34.3MB ± 0% 23.9MB ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="1",i=~"^.$",j!="foo" 51.6MB ± 0% 25.5MB ± 0% Querier/Block/labelValuesWithMatchers/n_with_j!="foo" 144MB ± 0% 139MB ± 0% Querier/Block/labelValuesWithMatchers/n_with_i="1" 6.43kB ± 0% 8.66kB ± 0% name old allocs/op new allocs/op Querier/Head/labelValuesWithMatchers/i_with_n="1" 104k ± 0% 500k ± 0% Querier/Head/labelValuesWithMatchers/i_with_n="^.+$" 204k ± 0% 600k ± 0% Querier/Head/labelValuesWithMatchers/i_with_n="1",j!="foo" 104k ± 0% 500k ± 0% Querier/Head/labelValuesWithMatchers/i_with_n="1",i=~"^.$",j!="foo" 204k ± 0% 500k ± 0% Querier/Head/labelValuesWithMatchers/n_with_j!="foo" 66.0 ± 0% 255.0 ± 0% Querier/Head/labelValuesWithMatchers/n_with_i="1" 61.0 ± 0% 205.0 ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="1" 304k ± 0% 600k ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="^.+$" 5.20M ± 0% 0.70M ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="1",j!="foo" 204k ± 0% 600k ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="1",i=~"^.$",j!="foo" 304k ± 0% 600k ± 0% Querier/Block/labelValuesWithMatchers/n_with_j!="foo" 3.00M ± 0% 0.00M ± 0% Querier/Block/labelValuesWithMatchers/n_with_i="1" 61.0 ± 0% 247.0 ± 0% Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> * Don't expand postings to intersect them Using a min heap we can check whether matched postings intersect with each one of the label values postings. This avoid expanding postings (and thus having all of them in memory at any point). Slightly slower than the expanding postings version for some cases, but definitely pays the price once the cardinality grows. Still offers 10x latency improvement where previous latencies were reaching 1s. Benchmark results: name \ time/op old.txt intersect.txt intersect_noexpand.txt Querier/Head/labelValuesWithMatchers/i_with_n="1" 157ms ± 0% 48ms ± 0% 110ms ± 0% Querier/Head/labelValuesWithMatchers/i_with_n="^.+$" 1.80s ± 0% 0.46s ± 0% 0.18s ± 0% Querier/Head/labelValuesWithMatchers/i_with_n="1",j!="foo" 144ms ± 0% 57ms ± 0% 125ms ± 0% Querier/Head/labelValuesWithMatchers/i_with_n="1",i=~"^.$",j!="foo" 304ms ± 0% 111ms ± 0% 177ms ± 0% Querier/Head/labelValuesWithMatchers/n_with_j!="foo" 761ms ± 0% 164ms ± 0% 134ms ± 0% Querier/Head/labelValuesWithMatchers/n_with_i="1" 6.11µs ± 0% 6.62µs ± 0% 4.29µs ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="1" 117ms ± 0% 62ms ± 0% 120ms ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="^.+$" 1.44s ± 0% 0.24s ± 0% 0.15s ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="1",j!="foo" 92.1ms ± 0% 70.3ms ± 0% 125.4ms ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="1",i=~"^.$",j!="foo" 196ms ± 0% 115ms ± 0% 170ms ± 0% Querier/Block/labelValuesWithMatchers/n_with_j!="foo" 1.23s ± 0% 0.21s ± 0% 0.14s ± 0% Querier/Block/labelValuesWithMatchers/n_with_i="1" 1.06ms ± 0% 0.88ms ± 0% 0.92ms ± 0% name \ alloc/op old.txt intersect.txt intersect_noexpand.txt Querier/Head/labelValuesWithMatchers/i_with_n="1" 29.5MB ± 0% 26.9MB ± 0% 19.1MB ± 0% Querier/Head/labelValuesWithMatchers/i_with_n="^.+$" 46.8MB ± 0% 251.5MB ± 0% 36.3MB ± 0% Querier/Head/labelValuesWithMatchers/i_with_n="1",j!="foo" 29.5MB ± 0% 22.3MB ± 0% 19.1MB ± 0% Querier/Head/labelValuesWithMatchers/i_with_n="1",i=~"^.$",j!="foo" 46.8MB ± 0% 23.9MB ± 0% 20.7MB ± 0% Querier/Head/labelValuesWithMatchers/n_with_j!="foo" 10.3kB ± 0% 138535.2kB ± 0% 6.4kB ± 0% Querier/Head/labelValuesWithMatchers/n_with_i="1" 5.54kB ± 0% 7.09kB ± 0% 4.30kB ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="1" 39.1MB ± 0% 28.5MB ± 0% 20.7MB ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="^.+$" 287MB ± 0% 253MB ± 0% 38MB ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="1",j!="foo" 34.3MB ± 0% 23.9MB ± 0% 20.7MB ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="1",i=~"^.$",j!="foo" 51.6MB ± 0% 25.5MB ± 0% 22.3MB ± 0% Querier/Block/labelValuesWithMatchers/n_with_j!="foo" 144MB ± 0% 139MB ± 0% 0MB ± 0% Querier/Block/labelValuesWithMatchers/n_with_i="1" 6.43kB ± 0% 8.66kB ± 0% 5.86kB ± 0% name \ allocs/op old.txt intersect.txt intersect_noexpand.txt Querier/Head/labelValuesWithMatchers/i_with_n="1" 104k ± 0% 500k ± 0% 300k ± 0% Querier/Head/labelValuesWithMatchers/i_with_n="^.+$" 204k ± 0% 600k ± 0% 400k ± 0% Querier/Head/labelValuesWithMatchers/i_with_n="1",j!="foo" 104k ± 0% 500k ± 0% 300k ± 0% Querier/Head/labelValuesWithMatchers/i_with_n="1",i=~"^.$",j!="foo" 204k ± 0% 500k ± 0% 300k ± 0% Querier/Head/labelValuesWithMatchers/n_with_j!="foo" 66.0 ± 0% 255.0 ± 0% 139.0 ± 0% Querier/Head/labelValuesWithMatchers/n_with_i="1" 61.0 ± 0% 205.0 ± 0% 87.0 ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="1" 304k ± 0% 600k ± 0% 400k ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="^.+$" 5.20M ± 0% 0.70M ± 0% 0.50M ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="1",j!="foo" 204k ± 0% 600k ± 0% 400k ± 0% Querier/Block/labelValuesWithMatchers/i_with_n="1",i=~"^.$",j!="foo" 304k ± 0% 600k ± 0% 400k ± 0% Querier/Block/labelValuesWithMatchers/n_with_j!="foo" 3.00M ± 0% 0.00M ± 0% 0.00M ± 0% Querier/Block/labelValuesWithMatchers/n_with_i="1" 61.0 ± 0% 247.0 ± 0% 129.0 ± 0% Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> * Apply comment suggestions from the code review Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> Co-authored-by: Ganesh Vernekar <15064823+codesome@users.noreply.github.com> * Change else { if } to else if Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> * Remove sorting of label values We were not sorting them before, so no need to sort them now Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> Co-authored-by: Ganesh Vernekar <15064823+codesome@users.noreply.github.com>	2021-12-28 15:59:03 +01:00
Ganesh Vernekar	6c7577177c	Merge pull request #10056 from charlesxsh/fix-TestWALRestoreCorrupted fix potential goroutine leaks at TestWALRestoreCorrupted	2021-12-21 14:54:45 +05:30
beorn7	86cc83b13c	storage: iterator fixes after merge Signed-off-by: beorn7 <beorn@grafana.com>	2021-12-18 14:12:01 +01:00
beorn7	64c7bd2b08	Merge branch 'main' into sparsehistogram	2021-12-18 14:04:25 +01:00
Shihao Xia	3696d7dedb	fix potential goroutine leaks Signed-off-by: Shihao Xia <charlesxsh@hotmail.com>	2021-12-17 18:35:30 -05:00
Julien Pivotto	27343277fa	Merge release-2.32 forward into main (#10032 ) * storage: expose bug in iterators #10027 Signed-off-by: beorn7 <beorn@grafana.com> * storage: fix bug #10027 in iterators' Seek method Signed-off-by: beorn7 <beorn@grafana.com> * Append reporting metrics without limit If reporting metrics fails due to reaching the limit, this makes the target appear as UP in the UI, but the metrics are missing. This commit bypasses that limit for report metrics. Signed-off-by: Julien Pivotto <roidelapluie@inuits.eu> * Remove check against cfg so interval/ timeout are always set (#10023) (#10031) Signed-off-by: Nicholas Blott <blottn@tcd.ie> Co-authored-by: Nicholas Blott <blottn@tcd.ie> * Cut v2.32.1 Signed-off-by: Julius Volz <julius.volz@gmail.com> * Apply suggestions from code review Signed-off-by: Julius Volz <julius.volz@gmail.com> Co-authored-by: Levi Harrison <git@leviharrison.dev> Co-authored-by: Julien Pivotto <roidelapluie@inuits.eu> Co-authored-by: Nicholas Blott <blottn@tcd.ie> Co-authored-by: Julius Volz <julius.volz@gmail.com> Co-authored-by: Levi Harrison <git@leviharrison.dev>	2021-12-17 23:18:38 +01:00
beorn7	0ede6ae321	storage: fix bug #10027 in iterators' Seek method Signed-off-by: beorn7 <beorn@grafana.com>	2021-12-16 12:07:35 +01:00
beorn7	6f33ab2b35	Merge branch 'main' into sparsehistogram	2021-12-15 13:49:33 +01:00
Ganesh Vernekar	227d155e3f	Fix queries after a failed snapshot replay (#9980 ) Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2021-12-09 15:15:27 +05:30
Ganesh Vernekar	05d4d97bcd	Fix queries after a failed snapshot replay (#9980 ) Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2021-12-08 15:32:14 +00:00
Nick Pillitteri	084bd70708	Convert atomic Int64 to native type when logging value (#9938 ) The atomic Int64 type wasn't able to be represented when logging via go-kit log and ended up as `"unsupported value type"`. Signed-off-by: Nick Pillitteri <nick.pillitteri@grafana.com>	2021-12-06 22:25:22 +01:00
beorn7	e8e9155a11	Merge branch 'main' into sparsehistogram	2021-11-30 18:22:37 +01:00
beorn7	e4e24453fa	Merge branch 'main' into beorn7/merge2	2021-11-30 17:19:06 +01:00
Robert Fratto	4cbddb41eb	tsdb/agent: Synchronize appender code with grafana/agent main (#9664 ) * tsdb/agent: synchronize appender code with grafana/agent main This commit synchronize the appender code with grafana/agent main. This includes adding support for appending exemplars. Closes #9610 Signed-off-by: Robert Fratto <robertfratto@gmail.com> * tsdb/agent: fix build error Signed-off-by: Robert Fratto <robertfratto@gmail.com> * tsdb/agent: introduce some exemplar tests, refactor tests a little Signed-off-by: Robert Fratto <robertfratto@gmail.com> * tsdb/agent: address review feedback - Re-use hash when creating a new series - Fix typo in exemplar append error Signed-off-by: Robert Fratto <robertfratto@gmail.com> * tsdb/agent: remove unused AddFast method Signed-off-by: Robert Fratto <robertfratto@gmail.com> * tsdb/agent: close wal reader after test Signed-off-by: Robert Fratto <robertfratto@gmail.com> * tsdb/agent: add out-of-order tracking, change series TS in commit Signed-off-by: Robert Fratto <robertfratto@gmail.com> * address review feedback Signed-off-by: Robert Fratto <robertfratto@gmail.com>	2021-11-30 21:14:40 +05:30
Björn Rabenstein	b866db009b	storage: Fix and improve the Seek method of various iterators (#9878 ) There was a subtle and nasty bug in listSeriesIterator.Seek. In addition, the Seek call is defined to be a no-op if the current position of the iterator is already pointing to a suitable sample. This commit adds fast paths for this case to several potentially expensive Seek calls. Another bug was in concreteSeriesIterator.Seek. It always searched the whole series and not from the current position of the iterator. Signed-off-by: beorn7 <beorn@grafana.com>	2021-11-29 15:17:56 +05:30
Björn Rabenstein	7e42acd3b1	tsdb: Rework iterators (#9877 ) - Pick At... method via return value of Next/Seek. - Do not clobber returned buckets. - Add partial FloatHistogram suppert. Note that the promql package is now _only_ dealing with FloatHistograms, following the idea that PromQL only knows float values. As a byproduct, I have removed the histogramSeries metric. In my understanding, series can have both float and histogram samples, so that metric doesn't make sense anymore. As another byproduct, I have converged the sampleBuf and the histogramSampleBuf in memSeries into one. The sample type stored in the sampleBuf has been extended to also contain histograms even before this commit. Signed-off-by: beorn7 <beorn@grafana.com>	2021-11-29 13:24:23 +05:30
Ganesh Vernekar	26c0a433f5	Support appending different sample types to the same series (#9705 ) * Support appending different sample types to the same series Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix comments Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix build Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2021-11-26 17:43:27 +05:30
Bryan Boreham	1b74a3812e	Fix panic, out of order chunks, and race warning during WAL replay (#9856 ) * Fix panic on WAL replay Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Refactor: introduce walSubsetProcessor walSubsetProcessor packages up the `processWALSamples()` function and its input and output channels, helping to clarify how these things relate. Signed-off-by: Bryan Boreham <bjboreham@gmail.com> * Refactor: extract more methods onto walSubsetProcessor This makes the main logic easier to follow. Signed-off-by: Bryan Boreham <bjboreham@gmail.com> * Fix race warning by locking processWALSamples Although we have waited for the processor to finish, we still get a warning from the race detector because it doesn't know how the different parts relate. Add a lock round each batch of samples, so the race detector can see that we never access series owned by the processor outside of a lock. Signed-off-by: Bryan Boreham <bjboreham@gmail.com> * Added test to reproduce issue 9859 Signed-off-by: Marco Pracucci <marco@pracucci.com> * Remove redundant unit test Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix out of order chunks during WAL replay Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix nits Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com> Co-authored-by: Marco Pracucci <marco@pracucci.com>	2021-11-25 13:36:14 +05:30
Oleg Zaytsev	5e746e4e88	Check postings bytes length when decoding (#9766 ) Added validation to expected postings length compared to the bytes slice length. With 32bit postings, we expect to have 4 bytes per each posting. If the number doesn't add up, we know that the input data is not compatible with our code (maybe it's cut, or padded with trash, or even written in a different coded). This is needed in downstream projects to correctly identify cached postings written with an unknown codec, but it's also a good idea to validate it here. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2021-11-24 15:26:37 +05:30
Darshan Chaudhary	9dcf8b2208	Add the ability to disable tsdb isolation (#9270 ) * Disable isolation in isolation struct Signed-off-by: darshanime <deathbullet@gmail.com> * Run tsdb tests with isolation disabled Signed-off-by: darshanime <deathbullet@gmail.com> * Check for isolation disabled in isoState.Close() Signed-off-by: darshanime <deathbullet@gmail.com> * use t.Skip to skip isolation tests when disabled Signed-off-by: darshanime <deathbullet@gmail.com> * address review comments Signed-off-by: darshanime <deathbullet@gmail.com> * fix test for defaultIsolationState Signed-off-by: darshanime <deathbullet@gmail.com> * Change flag name. Set flag in DB. Do not init txRing. Close isoState. Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Test disabled isolation in CircleCI test_go Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Skip isolation related tests in db_test.go Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com>	2021-11-19 15:41:32 +05:30
beorn7	5d4db805ac	Merge branch 'main' into sparsehistogram	2021-11-17 19:57:31 +01:00
Dieter Plaetinck	067efc3725	clarify HeadChunkID type and usage (#9726 ) Signed-off-by: Dieter Plaetinck <dieter@grafana.com> Co-authored-by: Ganesh Vernekar <15064823+codesome@users.noreply.github.com> Co-authored-by: Ganesh Vernekar <15064823+codesome@users.noreply.github.com>	2021-11-17 18:35:10 +05:30
Sunil Thaha	a484a83d4a	fix: panic when checkpoint directory is empty (#9687 ) Calling `wal.NewSegmentBufReader()` without any segments would cause a `panic` resulting in prometheus crashing. This patch fixes the panic by making segmentBufReader return a EOF if there are not segments. This also means an empty checkpoint directory which should never be the case unless it has been tampered with (or has issues due to the underlying filesystem e.g. NFS) would be ignored by Prometheus and would continue to run instead of the current behaviour which is to panic. Fixes: https://github.com/prometheus/prometheus/issues/9605 Signed-off-by: Sunil Thaha <sthaha@redhat.com>	2021-11-17 16:39:04 +05:30
Dieter Plaetinck	0fac9bb859	Add basic initial developer docs for TSDB (#9451 ) * Add basic initial developer docs for TSDB There's a decent amount of content already out there (blog posts, conference talks, etc), but: * when they get stale, they don't tend to get updated * they still leave me with questions that I'ld like to answer for developers (like me) who want to use, or work with, TSDB What I propose is developer docs inside the prometheus repository. Easy to find and harness the power of the community to expand it and keep it up to date. * perfect is the enemy of good. Let's have a base and incrementally improve * Markdown docs should be broad but not too deep. Source code comments can complement them, and are the ideal place for implementation details. Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * use example code that works out of the box Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * Apply suggestions from code review Co-authored-by: Ganesh Vernekar <15064823+codesome@users.noreply.github.com> Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * PR feedback Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * more docs Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * PR feedback Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * Apply suggestions from code review Signed-off-by: Dieter Plaetinck <dieter@grafana.com> Co-authored-by: Bartlomiej Plotka <bwplotka@gmail.com> * Apply suggestions from code review Signed-off-by: Dieter Plaetinck <dieter@grafana.com> Co-authored-by: Ganesh Vernekar <15064823+codesome@users.noreply.github.com> * feedback Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * Update tsdb/docs/usage.md Signed-off-by: Dieter Plaetinck <dieter@grafana.com> Co-authored-by: Ganesh Vernekar <15064823+codesome@users.noreply.github.com> * final tweaks Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * workaround docs versioning issue Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * Move example code to real executable, testable example. Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * cleanup example test and make sure it always reproduces Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * obtain temp dir in a way that works with older Go versions Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * Fix Ganesh's comments Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Co-authored-by: Ganesh Vernekar <15064823+codesome@users.noreply.github.com> Co-authored-by: Bartlomiej Plotka <bwplotka@gmail.com> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com>	2021-11-17 15:51:27 +05:30
beorn7	4c28d9fac7	Move to histogram.Histogram pointers This is to avoid copying the many fields of a histogram.Histogram all the time. This also fixes a bunch of formerly broken tests. Signed-off-by: beorn7 <beorn@grafana.com>	2021-11-12 23:17:35 +01:00
Mauro Stettler	8a4f659126	fix error message Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com>	2021-11-12 21:45:46 +01:00
Robert Fratto	72a9f7fee9	Share TSDB locker code with agent (#9623 ) * share tsdb db locker code with agent Closes #9616 Signed-off-by: Robert Fratto <robertfratto@gmail.com> * add flag to disable lockfile for agent Signed-off-by: Robert Fratto <robertfratto@gmail.com> * use agentOnlySetting instead of PreAction Signed-off-by: Robert Fratto <robertfratto@gmail.com> * tsdb: address review feedback 1. Rename Locker to DirLocker 2. Move DirLocker to tsdb/tsdbutil 3. Name metric using fmt.Sprintf 4. Refine error checking in DirLocker test Signed-off-by: Robert Fratto <robertfratto@gmail.com> * tsdb: create test utilities to assert expected DirLocker behavior Signed-off-by: Robert Fratto <robertfratto@gmail.com> * tsdb/tsdbutil: fix lint errors Signed-off-by: Robert Fratto <robertfratto@gmail.com> * tsdb/agent: fix windows test failure Use new DB variable instead of overriding the old one. Signed-off-by: Robert Fratto <robertfratto@gmail.com>	2021-11-11 11:45:25 -05:00
beorn7	f1065e44a4	model: String method for histogram.Histogram This includes a regular bucket iterator and a string method for histogram.Bucket. Signed-off-by: beorn7 <beorn@grafana.com>	2021-11-11 17:29:22 +01:00
Peter Štibraný	422e7839d4	Add more size checks when writing individual sections in the index. (#9710 ) * Add more size checks when writing individual sections in the index. Signed-off-by: Peter Štibraný <pstibrany@gmail.com> * Use uint and add comment about it. Signed-off-by: Peter Štibraný <pstibrany@gmail.com>	2021-11-11 15:44:28 +05:30
Mateusz Gozdek	83086aee00	tsdb/agent: use unique registry per tests So tests can run in parallel. Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>	2021-11-11 01:37:24 +01:00
Mateusz Gozdek	2f312ff4c5	tsdb: mark TestTombstoneCleanRetentionLimitsRace test as slow It takes over 100 seconds to execute this test, so I'd consider it as slow. Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>	2021-11-11 01:37:24 +01:00
Levi Harrison	7400e07fa9	Close DB in Agent tests (#9630 ) * Close agent db in tests Signed-off-by: Levi Harrison <git@leviharrison.dev> * Close first DB before opening second Signed-off-by: Levi Harrison <git@leviharrison.dev> * Use seperate variables for different DBs? Signed-off-by: Levi Harrison <git@leviharrison.dev> * Close remote storage Signed-off-by: Julien Pivotto <roidelapluie@inuits.eu> * Fix closing of stuff Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Remove the build flags after a rebase Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix closing of stuff 2 Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Co-authored-by: Julien Pivotto <roidelapluie@inuits.eu> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com>	2021-11-10 20:05:54 +05:30
Mateusz Gozdek	f4650c27e7	tsdb/wal: fix flaky TestReaderFuzz* tests It seems sometimes you can get error like: Error: Not equal: expected: []byte(nil) actual : []byte{} Diff: --- Expected +++ Actual @@ -1,2 +1,3 @@ -([]uint8) <nil> +([]uint8) { +} This commit does what bytes.Equal does to silence those differences. I'm not sure if this is a correct solution or just covering up the actual bug. Closes #9574 Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>	2021-11-09 14:32:20 +01:00
Björn Rabenstein	4c56a193c5	Merge pull request #9478 from prometheus/beorn7/pkg-deprecation Move packages out of deprecated pkg directory	2021-11-09 11:09:16 +01:00
曹明	a0d31c28fc	tsdb: Add windows arm64 support. Signed-off-by: 曹明 <caoming1@kingsoft.com>	2021-11-09 11:07:27 +01:00
Mateusz Gozdek	b319b14431	tsdb/chunks: preallocate at least some space on non-Windows systems (#9581 ) To avoid potential chunk corruption read, which I am not sure why is happening. Closes #9561. Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>	2021-11-09 13:47:00 +05:30
beorn7	c954cd9d1d	Move packages out of deprecated pkg directory This creates a new `model` directory and moves all data-model related packages over there: exemplar labels relabel rulefmt textparse timestamp value All the others are more or less utilities and have been moved to `util`: gate logging modetimevfs pool runtime Signed-off-by: beorn7 <beorn@grafana.com>	2021-11-09 08:03:10 +01:00
beorn7	a1e595edac	Fix two trivial lint warnings Not sure why those show up for me locally but not if run by the CI. Signed-off-by: beorn7 <beorn@grafana.com>	2021-11-08 22:32:13 +01:00
beorn7	8f92c90897	Add TODOs and some minor tweaks Signed-off-by: beorn7 <beorn@grafana.com>	2021-11-07 17:12:04 +01:00
Dieter Plaetinck	cda025b5b5	TSDB: demistify SeriesRefs and ChunkRefs (#9536 ) * TSDB: demistify seriesRefs and ChunkRefs The TSDB package contains many types of series and chunk references, all shrouded in uint types. Often the same uint value may actually mean one of different types, in non-obvious ways. This PR aims to clarify the code and help navigating to relevant docs, usage, etc much quicker. Concretely: * Use appropriately named types and document their semantics and relations. * Make multiplexing and demuxing of types explicit (on the boundaries between concrete implementations and generic interfaces). * Casting between different types should be free. None of the changes should have any impact on how the code runs. TODO: Implement BlockSeriesRef where appropriate (for a future PR) Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * feedback Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * agent: demistify seriesRefs and ChunkRefs Signed-off-by: Dieter Plaetinck <dieter@grafana.com>	2021-11-06 15:40:04 +05:30
johncming	b882d2b7c7	tsdb/wal: Avoid writing closed channel. (#9566 ) Signed-off-by: johncming <johncming@yahoo.com>	2021-11-06 15:11:06 +05:30
chenlujjj	d18e42c650	refine comments of Checkpoint function (#9655 ) Signed-off-by: chenlujjj <953546398@qq.com>	2021-11-06 15:09:16 +05:30
Marco Pracucci	309b094b92	Optimized MemPostings.EnsureOrder() (#9673 ) * Optimizes MemPostings.EnsureOrder() Signed-off-by: Marco Pracucci <marco@pracucci.com> * Ignore linter warning Signed-off-by: Marco Pracucci <marco@pracucci.com>	2021-11-05 10:01:23 +00:00
Ganesh Vernekar	c8b267efd6	Get histograms from TSDB to the rate() function implementation Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2021-11-03 19:04:18 +05:30
Marco Pracucci	9f5ff5b269	Allow to disable trimming when querying TSDB (#9647 ) * Allow to disable trimming when querying TSDB Signed-off-by: Marco Pracucci <marco@pracucci.com> * Addressed review comments Signed-off-by: Marco Pracucci <marco@pracucci.com> * Added unit test Signed-off-by: Marco Pracucci <marco@pracucci.com> * Renamed TrimDisabled to DisableTrimming Signed-off-by: Marco Pracucci <marco@pracucci.com>	2021-11-03 15:38:34 +05:30
Marco Pracucci	edd05d7010	Add Head.AppendableMinValidTime() (#9643 ) Signed-off-by: Marco Pracucci <marco@pracucci.com>	2021-11-03 13:09:54 +05:30
Mateusz Gozdek	b7bdf6fab2	Fix imports formatting According to `2829908806 (r58457095)`. Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>	2021-11-02 19:52:34 +01:00
Mateusz Gozdek	1a6c2283a3	Format Go source files using 'gofumpt -w -s -extra' Part of #9557 Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>	2021-11-02 19:52:34 +01:00
Julien Pivotto	6e1d6edb33	Exclude agent from windows tests (#9645 ) We are aware of the issue, but while we are working on it, having main tests broken is an annoyance. Signed-off-by: Julien Pivotto <roidelapluie@inuits.eu>	2021-11-02 13:58:51 +01:00
chenlujjj	660329d5b3	add tombstoneFormatVersionSize & tombstonesCRCSize constants (#9625 ) Signed-off-by: chenlujjj <953546398@qq.com>	2021-11-01 16:05:19 +05:30
Praveen Ghuge	64d9b41998	Use testing.T.TempDir() instead of ioutil.TempDir() in tsdb/wal unit tests (#9602 ) Signed-off-by: Praveen Ghuge <praveen.ghuge@outlook.com>	2021-11-01 12:28:18 +05:30
Robert Fratto	bc72a718c4	Initial draft of prometheus-agent (#8785 ) * Initial draft of prometheus-agent This commit introduces a new binary, prometheus-agent, based on the Grafana Agent code. It runs a WAL-only version of prometheus without the TSDB, alerting, or rule evaluations. It is intended to be used to remote_write to Prometheus or another remote_write receiver. By default, prometheus-agent will listen on port 9095 to not collide with the prometheus default of 9090. Truncation of the WAL cooperates on a best-effort case with Remote Write. Every time the WAL is truncated, the minimum timestamp of data to truncate is determined by the lowest sent timestamp of all samples across all remote_write endpoints. This gives loose guarantees that data from the WAL will not try to be removed until the maximum sample lifetime passes or remote_write starts functionining. Signed-off-by: Robert Fratto <robertfratto@gmail.com> * add tests for Prometheus agent (#22) * add tests for Prometheus agent * add tests for Prometheus agent * rearranged tests as per the review comments * update tests for Agent * changes as per code review comments Signed-off-by: SriKrishna Paparaju <paparaju@gmail.com> * incremental changes to prometheus agent Signed-off-by: SriKrishna Paparaju <paparaju@gmail.com> * changes as per code review comments Signed-off-by: SriKrishna Paparaju <paparaju@gmail.com> * Commit feedback from code review Co-authored-by: Bartlomiej Plotka <bwplotka@gmail.com> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com> Signed-off-by: Robert Fratto <robertfratto@gmail.com> * Port over some comments from grafana/agent Signed-off-by: Robert Fratto <robertfratto@gmail.com> * Rename agent.Storage to agent.DB for tsdb consistency Signed-off-by: Robert Fratto <robertfratto@gmail.com> * Consolidate agentMode ifs in cmd/prometheus/main.go Signed-off-by: Robert Fratto <robertfratto@gmail.com> * Document PreAction usage requirements better for agent mode flags Signed-off-by: Robert Fratto <robertfratto@gmail.com> * remove unnecessary defaultListenAddr Signed-off-by: Robert Fratto <robertfratto@gmail.com> * `go fmt ./tsdb/agent` and fix lint errors Signed-off-by: Robert Fratto <robertfratto@gmail.com> Co-authored-by: SriKrishna Paparaju <paparaju@gmail.com>	2021-10-29 16:25:05 +01:00
Xiaochao Dong	c2d1c85857	close tsdb.head in test case (#9580 ) Signed-off-by: Xiaochao Dong (@damnever) <the.xcdong@gmail.com>	2021-10-26 11:36:25 +05:30
Furkan Türkal	0c07663b70	fix: possible race on shared variables in test (#9470 ) Fixes #9433 Signed-off-by: Furkan <furkan.turkal@trendyol.com>	2021-10-25 18:44:40 +05:30
Dieter Plaetinck	d5bfbe3114	improve bstream comments and doc (#9560 ) * improve bstream comments and doc Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * feedback Signed-off-by: Dieter Plaetinck <dieter@grafana.com>	2021-10-25 18:44:15 +05:30
Julien Pivotto	73255e15f6	Address golint failures from revive Signed-off-by: Julien Pivotto <roidelapluie@inuits.eu>	2021-10-23 00:53:11 +02:00
Serge Catudal	8c3eca84db	Fix remote write receiver endpoint for exemplars (#9414 ) Signed-off-by: Serge Catudal <serge.catudal@gmail.com>	2021-10-21 22:58:40 +02:00
beorn7	a9008f5423	Merge branch 'main' into sparsehistogram	2021-10-19 17:14:23 +02:00
beorn7	4998b9750f	chunkenc: Bugfix and naming tweaks Signed-off-by: beorn7 <beorn@grafana.com>	2021-10-19 15:38:32 +02:00
beorn7	78ef9c6359	chunkenc: make xor reading more DRY Signed-off-by: beorn7 <beorn@grafana.com>	2021-10-19 15:28:33 +02:00
beorn7	4a1b84f8b2	chunkenc: make xor writing more DRY Signed-off-by: beorn7 <beorn@grafana.com>	2021-10-19 15:28:33 +02:00
Björn Rabenstein	3704c6c20a	Merge pull request #9533 from prometheus/beorn7/sparsehistogram tsdb: Complete chunk format documentation	2021-10-19 13:51:46 +02:00
beorn7	1a4e54cfbb	tsdb: Complete chunk format documentation This also tweaks and fixes a few things done previously. Signed-off-by: beorn7 <beorn@grafana.com>	2021-10-19 13:51:30 +02:00
beorn7	0876d57aea	chunkenc: Add test for chunk layout encoding And fix a bug exposed by it... Signed-off-by: beorn7 <beorn@grafana.com>	2021-10-18 19:37:24 +02:00
beorn7	ad9b4c2b68	Fix typos Signed-off-by: beorn7 <beorn@grafana.com>	2021-10-18 15:44:13 +02:00
beorn7	fe50d6fc14	Update chunk layout documentation Signed-off-by: beorn7 <beorn@grafana.com>	2021-10-15 23:18:41 +02:00
beorn7	ed33aea392	Avoid redundant varint decoding in chunk appender construction Signed-off-by: beorn7 <beorn@grafana.com>	2021-10-15 20:33:14 +02:00
beorn7	d31bb75dc4	Use VarbitUint rather than VarbitInt to encode len(spans) Signed-off-by: beorn7 <beorn@grafana.com>	2021-10-15 15:27:32 +02:00
beorn7	3179215a59	Encode zero threshold first This guaranees that the zero threshold is byte-aligned. Not sure if that helps in any way, but at least it won't harm. Signed-off-by: beorn7 <beorn@grafana.com>	2021-10-14 14:55:21 +02:00
beorn7	c5522677bf	Improve encoding of zero threshold Signed-off-by: beorn7 <beorn@grafana.com>	2021-10-14 14:47:26 +02:00
beorn7	7093b089f2	Use more varbit in histogram chunks This adds bit buckets for larger numbers to varbit encoding and also an unsigned version of varbit encoding. Then, varbit encoding is used for all the histogram chunk data instead of varint. Signed-off-by: beorn7 <beorn@grafana.com>	2021-10-13 20:03:35 +02:00
Björn Rabenstein	7309c20e7e	Merge pull request #9500 from codesome/resettests Add unit test for counter reset header	2021-10-13 18:19:21 +02:00
Ganesh Vernekar	dcaf568279	Metadata -> Layout renaming Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2021-10-13 20:27:48 +05:30
Ganesh Vernekar	4e206c7c77	Fix reviews Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2021-10-13 20:23:31 +05:30
Dieter Plaetinck	d5afe0a577	TSDB: Use a dedicated head chunk reference type (#9501 ) * Use dedicated Ref type Throughout the code base, there are reference types masked as regular integers. Let's use dedicated types. They are equivalent, but clearer semantically. This also makes it trivial to find where they are used, and from uses, find the centralized docs. Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * postpone some work until after possible return Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * clarify Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * rename feedback Signed-off-by: Dieter Plaetinck <dieter@grafana.com> * skip header is up to caller Signed-off-by: Dieter Plaetinck <dieter@grafana.com>	2021-10-13 17:44:32 +05:30

... 6 7 8 9 10 ...

1173 commits