prometheus

mirror of https://github.com/prometheus/prometheus.git synced 2024-12-24 05:04:05 -08:00

Author	SHA1	Message	Date
Ganesh Vernekar	c6f3d4ab33	Remove temporary patch for out-of-order (#283 ) * Remove temporary patch for out-of-order Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Remove ooo_wbl patch and fix tests Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-07-04 13:35:18 +00:00
Ganesh Vernekar	5e8406a1d4	Avoid gaps in in-order data after restart with out-of-order enabled (#277 ) * Avoid gaps in in-order data after restart with out-of-order enabled Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix tests, do the temporary patch only if OOO is enabled Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Avoid Peter's confusion Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Use latest OutOfOrderTimeWindow Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-06-27 20:26:25 +05:30
Jesus Vazquez	1446b53d87	Merge pull request #276 from grafana/jvp/rename-oooallowance-to-oootimewindow Rename OutOfOrderAllowance to OutOfOrderTimeWindow	2022-06-24 12:40:20 +02:00
Ganesh Vernekar	abde1e0ba1	Update MinOOOTime and MaxOOOTime properly after restart (#275 ) Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-06-24 10:26:33 +00:00
Jesus Vazquez	e70e769889	Rename OutOfOrderAllowance to OutOfOrderTimeWindow After review Allowance is perhaps a bit misleading so we've decided to replace it with a more common term like TimeWindow.	2022-06-24 12:23:38 +02:00
Ganesh Vernekar	df59320886	Add out-of-order sample support to the TSDB (#269 ) This implementation is based on this design doc: https://docs.google.com/document/d/1Kppm7qL9C-BJB1j6yb6-9ObG3AbdZnFUBYPNNWwDBYM/edit?usp=sharing This commit adds support to accept out-of-order ("OOO") sample into the TSDB up to a configurable time allowance. If OOO is enabled, overlapping querying are automatically enabled. Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> Co-authored-by: Jesus Vazquez <jesus.vazquez@grafana.com> Co-authored-by: Ganesh Vernekar <ganeshvern@gmail.com> Co-authored-by: Dieter Plaetinck <dieter@grafana.com>	2022-06-22 11:45:21 +00:00
Peter Štibraný	cf7aeb59a7	Merge remote-tracking branch 'upstream/main' into update-prometheus	2022-06-14 09:34:59 +02:00
Jesus Vazquez	06f1d3c349	Merge pull request #251 from grafana/codesome/ooopatch Add an option to enable overlapping compaction separately with overlapping queries	2022-06-13 17:11:38 +02:00
Bryan Boreham	542b9ecdbd	tsdb: reduce sleep time when reading WAL (#10859 ) The code sleeps for a short time to allow goroutines to finish, however it seems the duration can be reduced a lot, speeding up the reading process. I checked using some WAL data from production, and the queue is almost always empty at the time we enter `waitForIdle()` so there is no danger of spinning in the tight loop. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-06-12 11:54:11 +05:30
Ganesh Vernekar	0eb828c179	Add an option to enable overlapping compaction separately with overlapping queries Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-06-09 12:11:42 -07:00
Peter Štibraný	d051065441	Remove use of io/ioutil.	2022-06-09 15:01:34 +02:00
Peter Štibraný	9d51bf50db	Merge upstream Prometheus	2022-06-09 11:29:19 +02:00
Peter Štibraný	55236be04a	Fix comments. (#248 )	2022-06-08 10:01:05 +02:00
Bryan Boreham	9f79a6f4b5	tsdb: faster CRC check by avoiding allocations (#10789 ) Instead of creating a new hashing object every time, call `crc32.Checksum` which computes the answer without allocations. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-06-08 08:00:59 +05:30
Peter Štibraný	1e2d2fb2d8	Job queue (#247 ) This PR reimplements chan chunkWriteJob with custom buffered queue that should use less memory, because it doesn't preallocate entire buffer for maximum queue size at once. Instead it allocates individual "segments" with smaller size. As elements are added to the queue, they fill individual segments. When elements are removed from the queue (and segments), empty segments can be thrown away. This doesn't change memory usage of the queue when it's full, but should decrease its memory footprint when it's empty (queue will keep max 1 segment in such case).	2022-06-07 17:42:28 +02:00
Mauro Stettler	459f59935c	Reduce chunk write queue memory usage (#131 ) * dont waste space on the chunkRefMap * add time factor * add comments * better readability * add instrumentation and more comments * formatting * uppercase comments * Address review feedback. Renamed "free" to "shrink" everywhere, updated comments and threshold to 1000. * double space Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> Co-authored-by: Peter Štibraný <pstibrany@gmail.com>	2022-06-02 13:48:30 +02:00
Matej Gera	1dd247f68b	Remote Write: Rename confusing `walDir` parameter to `dir` (#10464 ) * Rename walDir parameter to dir Signed-off-by: Matej Gera <matejgera@gmail.com> * Improve NewQueueManager comment Signed-off-by: Matej Gera <matejgera@gmail.com>	2022-05-30 21:45:30 -07:00
David Leadbeater	57f4aab27d	Update godoc links and remove note about TSDB versioning (#10754 ) Signed-off-by: David Leadbeater <dgl@dgl.cx>	2022-05-26 18:34:43 +10:00
maizige	10b677b826	fix typo (#10696 ) Update doc comment Signed-off-by: gemaizi <864321211@qq.com>	2022-05-25 18:01:45 +02:00
Filip Petkovski	d3cb39044e	Fix typo in symbol table size exceeded error message (#10746 ) This commit fixes a typo when reporting an error that the the symbols table size has been exceeded. Signed-off-by: Filip Petkovski <filip.petkovsky@gmail.com>	2022-05-25 10:40:36 +02:00
Julien Pivotto	6e3a0efe40	Make necessary change to compile promql parser to wasm (#10683 ) Signed-off-by: Julien Pivotto <roidelapluie@o11y.eu>	2022-05-12 09:12:05 +02:00
Matthias Rampke	78f2645787	test(tsdb): break up repeated test to avoid timeout (#10671 ) On macOS, the TestTombstoneCleanRetentionLimitsRace performs very poorly. It takes more than a second to write out one block, and as it writes 400 of them, we run into the 10-minute test timeout frequently. While this doesn't fix the actual performance issue, breaking each iteration into a subtest makes the test pass reliably (because each iteration comfortably finishes in under a minute). Related report: https://groups.google.com/g/prometheus-developers/c/jxQ6Ayg6VJ4/m/03H_DS9PDAAJ Signed-off-by: Matthias Rampke <matthias@prometheus.io>	2022-05-09 00:39:26 +02:00
Łukasz Mierzwa	88f9b248b4	Correctly format error message (#10669 ) Signed-off-by: Łukasz Mierzwa <l.mierzwa@gmail.com>	2022-05-06 00:42:31 +02:00
Bryan Boreham	4b9f248e85	unit tests: make all Labels sorted alphabetically (#10532 ) "Labels is a sorted set of labels. Order has to be guaranteed upon instantiation." says the comment, so fix all the tests that break this rule. For `BenchmarkLabelValuesWithMatchers()` and `BenchmarkHeadLabelValuesWithMatchers()` the amount of work done changes significantly if you put the labels in order, because all series refs get neatly partitioned by the `tens` label, so I renamed the labels to maintain the previous behaviour. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-05-04 23:41:36 +02:00
Matthieu MOREL	e2ede285a2	refactor: move from io/ioutil to io and os packages (#10528 ) * refactor: move from io/ioutil to io and os packages * use fs.DirEntry instead of os.FileInfo after os.ReadDir Signed-off-by: MOREL Matthieu <matthieu.morel@cnp.fr>	2022-04-27 11:24:36 +02:00
Mauro Stettler	64e6c171c2	Merge pull request #216 from grafana/merge_upstream add option to use the new chunk disk mapper from upstream	2022-04-25 11:27:15 -04:00
Mauro Stettler	55cbbafe38	update comment Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com>	2022-04-25 15:05:21 +00:00
Oleg Zaytsev	9d66af50a8	Fix TestMemSeries_append_atVariableRate Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2022-04-20 17:50:05 +02:00
Oleg Zaytsev	ac87e2d4d6	Merge remote-tracking branch 'prometheus/main' into update-prometheus	2022-04-20 17:39:51 +02:00
Oleg Zaytsev	af0f6da5cb	Fix chunk overflow appending samples at a variable rate (#10607 ) * Add a test with variable samples rate append This test overflows the chunk created in memseries, and the total amount of samples in the (only) mmapped chunk is 29, instead of the 65565 appended ones. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> * Cut new chunk when rate prediction was wrong When appending samples at a slow rate, and then appending at a higher rate, the prediction we made to cut a new chunk is no longer valid. Sometimes this can even cause an overflow in the chunk, if more samples than uint16 can hold are appended. Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> * Improve comment on 2samplesPerChunk Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com> Assert that all chunks have less than 240 samples Also, trigger new chunk at 240, not at more than 240 Signed-off-by: Oleg Zaytsev <mail@olegzaytsev.com>	2022-04-20 14:54:20 +02:00
Mauro Stettler	00f1b1556c	undo renaming Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com>	2022-04-19 20:16:20 +00:00
Mauro Stettler	2eef3e76b8	check return status of cutnewfile Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com>	2022-04-19 20:15:37 +00:00
Mauro Stettler	3fa636a799	fix unit test Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com>	2022-04-19 19:57:33 +00:00
Mauro Stettler	04114aa2e6	add old chunk disk mapper back Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com>	2022-04-19 19:21:22 +00:00
Mauro Stettler	bd50b04fed	remove old chunk disk mapper initialization and replace it with the upstream one Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com>	2022-04-18 18:03:05 +00:00
Paschalis Tsilias	40c1efe8bc	tsdb/agent: Ignore duplicate exemplars (#10595 ) * tsdb/agent: Ignore duplicate exemplars Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Make each exemplar unique in TestCommit Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Re-Trigger CI for Windows and UI-related steps Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Change test comment to properly re-trigger pipeline Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com> * Defer Close() calls for test agent and segment reader Signed-off-by: Paschalis Tsilias <paschalist0@gmail.com>	2022-04-18 11:41:04 -04:00
Jesus Vazquez	7106db9303	Merge remote-tracking branch 'upstream/main' into jvp/merge-prometheus-main	2022-04-15 12:09:46 +02:00
Julien Pivotto	685ce9964d	Merge pull request #10599 from prometheus/release-2.35 Merge back release 2.35	2022-04-15 00:10:06 +02:00
Robert Fratto	286dfc70b7	tsdb/agent: port grafana/agent#676 (#10587 ) * tsdb/agent: port grafana/agent#676 grafana/agent#676 fixed an issue where a loading a WAL with multiple segments may result in ref ID collision. The starting ref ID for new series should be set to the highest ref ID across all series records from all WAL segments. This fixes an issue where the starting ref ID was incorrectly set to the highest ref ID found in the newest segment, which may not have any ref IDs at all if no series records have been appended to it yet. Signed-off-by: Robert Fratto <robertfratto@gmail.com> * tsdb/agent: update terminology (s/ref ID/nextRef) Signed-off-by: Robert Fratto <robertfratto@gmail.com>	2022-04-14 10:27:06 +02:00
Jesus Vazquez	51f26e0cba	Necessary changes to make the merge work This commit disables some unused workflows on our CI. Also uses grafana/regexp instead of regexp which is blackisted. Also updates head_test TestHeadReadWriterRepair increasing ChunkWriteQueueSize to 1 so that the chunk disk mapper uses the async queue. This seems to be default behaviour in upstream prometheus and without this option our test fails.	2022-04-13 14:45:43 +02:00
Jesus Vazquez	48aa5cd096	Merge remote-tracking branch 'upstream/main' into jvp/merge-prometheus-main	2022-04-12 16:40:00 +02:00
Jesus Vazquez	c02b13b7f4	Discard unknown chunk encodings (#196 ) * Chunks replay skips chunks with unknown encodings We've changed the logic of loadMmappedChunks to skip chunks that have unknown encodings. To do so we've modified IterateAllChunks to accept an extra encoding argument in the callback function. Also added unit tests in the head and chunk disk mapper. * Also add an unit test for the old chunk diskmapper * s/createUnsupportedChunk/writeUnsupportedChunk/g	2022-04-12 10:35:10 +00:00
chavacava	0b41fd6e71	Fix data races in WAL replay (#10571 ) Signed-off-by: chavacava <salvadorcavadini+github@gmail.com>	2022-04-12 16:00:20 +05:30
Bryan Boreham	2c1be4df7b	tsdb: more efficient sorting of postings read from WAL at startup (#10500 ) * tsdb: avoid slice-to-interface allocation in EnsureOrder This is pulling the `seriesRefSlice` out of the loop, so the compiler doesn't allocate a new one on the heap every time. Signed-off-by: Bryan Boreham <bjboreham@gmail.com> * tsdb: use pointer type in Pool for EnsureOrder As noted by staticcheck, Pool prefers the objects in the pool to have pointer type. This is a little more fiddly to code, but avoids allocation of a wrapper object every time a slice is put into the pool. Removed a comment that said fixing this has a performance penalty: not borne out by benchmarks. Signed-off-by: Bryan Boreham <bjboreham@gmail.com>	2022-03-30 15:10:19 +05:30
Wilbert Guo	83a2e52bc2	Add SyncForState Implementation for Ruler HA (#10070 ) * continuously syncing activeAt for alerts Signed-off-by: Yijie Qin <qinyijie@amazon.com> Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> * add import Signed-off-by: Yijie Qin <qinyijie@amazon.com> Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> * Refactor SyncForState and add unit tests Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> * Format code Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> * Add hook for syncForState Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> Fix go lint Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> Refactor syncForState override implementation Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> Add syncForState override func as argument to Update() Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> Fix go formatting Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> Fix circleci test errors Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> Remove overrideFunc as argument to run() Signed-off-by: Wilbert Guo <wilbeguo@amazon.com> * remove the syncForState Signed-off-by: Yijie Qin <qinyijie@amazon.com> * use the override function to decide if need to replace the activeAt or not Signed-off-by: Yijie Qin <qinyijie@amazon.com> * fix test case Signed-off-by: Yijie Qin <qinyijie@amazon.com> * fix format Signed-off-by: Yijie Qin <qinyijie@amazon.com> * Trigger build Signed-off-by: Yijie Qin <qinyijie@amazon.com> * fixing comments Signed-off-by: Yijie Qin <qinyijie@amazon.com> * return the result of map of alerts instead of single one Signed-off-by: Yijie Qin <qinyijie@amazon.com> * upper case the QueryforStateSeries Signed-off-by: Yijie Qin <qinyijie@amazon.com> * use a more generic rule group post process function type Signed-off-by: Yijie Qin <qinyijie@amazon.com> * fix indentation Signed-off-by: Yijie Qin <qinyijie@amazon.com> * fix gofmt Signed-off-by: Yijie Qin <qinyijie@amazon.com> * fix lint Signed-off-by: Yijie Qin <qinyijie@amazon.com> * fixing naming Signed-off-by: Yijie Qin <qinyijie@amazon.com> * fix comments Signed-off-by: Yijie Qin <qinyijie@amazon.com> * add the lastEvalTimestamp as parameter Signed-off-by: Yijie Qin <qinyijie@amazon.com> * fmt Signed-off-by: Yijie Qin <qinyijie@amazon.com> * change funcType to func Signed-off-by: Yijie Qin <qinyijie@amazon.com> Co-authored-by: Yijie Qin <qinyijie@amazon.com> Co-authored-by: Yijie Qin <63399121+qinxx108@users.noreply.github.com>	2022-03-29 02:16:46 +02:00
Howie	1291ec7185	deleting .tmp WAL files on startup (#10317 ) fix issue #10245 Signed-off-by: lihaowei <haoweili35@gmail.com> * minor changes Signed-off-by: lihaowei <haoweili35@gmail.com> * review changes Signed-off-by: lihaowei <haoweili35@gmail.com> * minor changes Signed-off-by: lihaowei <haoweili35@gmail.com>	2022-03-24 16:14:14 +05:30
Chris Marchbanks	c1387494dd	Merge pull request #10452 from prometheus/release-2.34 Merge Release 2.34 into main	2022-03-15 12:32:18 -06:00
songjiayang	8d8be43824	Update wal.md (#10442 ) update exemplar record type Signed-off-by: songjiayang <songjiayang1@gmail.com>	2022-03-15 22:33:45 +05:30
Ganesh Vernekar	23ce9ad9f0	Introduce evaluation delay for rule groups (#155 ) * Allow having evaluation delay for rule groups Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix lint Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Move the option to ManagerOptions Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Include evaluation_delay in the group config Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com> * Fix comments Signed-off-by: Ganesh Vernekar <ganeshvern@gmail.com>	2022-03-14 13:20:07 +00:00
Mauro Stettler	b025390cb4	Disable chunk write queue by default, allow user to configure the exact size (#10425 ) * Disable chunk write queue by default Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com> * update flag description Signed-off-by: Mauro Stettler <mauro.stettler@gmail.com>	2022-03-11 17:26:59 +01:00

1 2 3 4 5 ...

553 commits