restic

mirror of https://github.com/octoleo/restic.git synced 2024-11-22 21:05:10 +00:00

Author	SHA1	Message	Date
Michael Eischer	1ebd57247a	repository: optimize MasterIndex.Each Sending data through a channel at very high frequency is extremely inefficient. Thus use simple callbacks instead of channels. > name old time/op new time/op delta > MasterIndexEach-16 6.68s ±24% 0.96s ± 2% -85.64% (p=0.008 n=5+5)	2022-09-24 12:21:59 +02:00
Michael Eischer	0a6fa602c8	add option for setting min pack size	2022-08-05 23:47:12 +02:00
Michael Eischer	04e49924fb	checker: Fix S3 legacy layout detection	2022-07-23 11:19:32 +02:00
Michael Eischer	fcb3ddf181	check: Complain about usage of s3 legacy layout	2022-07-23 11:19:32 +02:00
Michael Eischer	8b8bd4e8ac	check: complain about mixed pack files	2022-07-23 11:19:32 +02:00
Michael Eischer	89d3ce852b	repository: extract Load/StoreJSONUnpacked A Load/Store method for each data type is much clearer. As a result the repository no longer needs a method to load / store json.	2022-07-17 13:22:00 +02:00
Michael Eischer	fbcbd5318c	repository: extract LoadTree/SaveTree The repository has no real idea what a Tree is. So these methods never belonged there.	2022-07-17 13:11:28 +02:00
Michael Eischer	6f53ecc1ae	adapt workers based on whether an operation is CPU or IO-bound Use runtime.GOMAXPROCS(0) as worker count for CPU-bound tasks, repo.Connections() for IO-bound task and a combination if a task can be both. Streaming packs is treated as IO-bound as adding more worker cannot provide a speedup. Typical IO-bound tasks are download / uploading / deleting files. Decoding / Encoding / Verifying are usually CPU-bound. Several tasks are a combination of both, e.g. for combined download and decode functions. In the latter case add both limits together. As the backends have their own concurrency limits restic still won't download more than repo.Connections() files in parallel, but the additional workers can decode already downloaded data in parallel.	2022-07-03 12:19:26 +02:00
Michael Eischer	120ccc8754	repository: Rework blob saving to use an async pack uploader Previously, SaveAndEncrypt would assemble blobs into packs and either return immediately if the pack is not yet full or upload the pack file otherwise. The upload will block the current goroutine until it finishes. Now, the upload is done using separate goroutines. This requires changes to the error handling. As uploads are no longer tied to a SaveAndEncrypt call, failed uploads are signaled using an errgroup. To count the uploaded amount of data, the pack header overhead is no longer returned by `packer.Finalize` but rather by `packer.HeaderOverhead`. This helper method is necessary to continue returning the pack header overhead directly to the responsible call to `repository.SaveBlob`. Without the method this would not be possible, as packs are finalized asynchronously.	2022-07-02 22:42:34 +02:00
Michael Eischer	5e0f1c3cef	check: remove dead code	2022-07-02 19:28:57 +02:00
Michael Eischer	0df022fa6d	check: Print full ids The short ids are not always unique. In addition, recovering from damages is easier when having the full ids as that makes it easier to access the corresponding files.	2022-07-02 19:28:57 +02:00
Alexander Neumann	99634c0936	Return real size from SaveBlob	2022-07-02 18:55:12 +02:00
MichaelEischer	fdc53a9d32	Merge pull request #3787 from MichaelEischer/refactor-repository repository: (Mostly) index-related cleanups	2022-07-02 18:54:04 +02:00
Michael Eischer	a77d5c4d11	repository: index saving belongs into the MasterIndex	2022-07-02 18:38:56 +02:00
greatroar	a0fa9c6e9f	Revert "restic prune: Merge three loops over the index" This reverts commit `8bdfcf779f`. Should fix #3809. Also needed to make #3290 apply cleanly.	2022-06-30 15:27:34 +02:00
MichaelEischer	19581dbc18	Merge pull request #3786 from greatroar/prune restic prune: Merge three loops over the index	2022-06-18 16:54:50 +02:00
greatroar	8bdfcf779f	restic prune: Merge three loops over the index There were three loops over the index in restic prune, to find duplicates, to determine sizes (in pack.Size) and to generate packInfos. These three are now one loop. This way, prune doesn't need to construct a set of duplicate blobs, pack.Size doesn't need to contain special logic for prune's use case (the onlyHdr argument) and pack.Size doesn't need to construct a map only to have it immediately transformed into a different map. Some quick testing on a 160GiB local repo doesn't show running time or memory use of restic prune --dry-run changing significantly.	2022-06-18 10:40:33 +02:00
greatroar	f92ecf13c9	all: Move away from pkg/errors, easy cases github.com/pkg/errors is no longer getting updates, because Go 1.13 went with the more flexible errors.{As,Is} function. Use those instead: errors from pkg/errors already support the Unwrap interface used by 1.13 error handling. Also: * check for io.EOF with a straight ==. That value should not be wrapped, and the chunker (whose error is checked in the cases changed) does not wrap it. * Give custom Error methods pointer receivers, so there's no ambiguity when type-switching since the value type will no longer implement error. * Make restic.ErrAlreadyLocked private, and rename it to alreadyLockedError to match the stdlib convention that error type names end in Error. * Same with rest.ErrIsNotExist => rest.notExistError. * Make s3.Backend.IsAccessDenied a private function.	2022-06-14 08:36:38 +02:00
Michael Eischer	5815f727ee	checker: convert error type to use pointer-receivers	2022-05-09 22:31:30 +02:00
Alexander Neumann	8b11b86383	Add option global --compression	2022-04-30 11:34:10 +02:00
Michael Eischer	7b9ae91e04	copy: Load snapshots before indexes	2022-04-09 12:27:25 +02:00
Michael Eischer	3d29083e60	copy/find/ls/recover/stats: Memorize snapshot listing before index These commands filter the snapshots according to some criteria which essentially requires loading the index before filtering the snapshots. Thus create a copy of the snapshots list beforehand and use it later on.	2022-04-09 12:26:30 +02:00
Michael Eischer	a773cb6527	pack: cleanup header size calculation	2022-03-28 22:09:49 +02:00
Michael Eischer	6408686973	repository: Simplify Blob equality check	2022-03-28 22:09:49 +02:00
Michael Eischer	f78bd14e28	repository: Remove pack implementation details from MasterIndex	2022-03-28 22:09:49 +02:00
Michael Eischer	4b3dc415ef	checker: cleanup header extraction	2022-02-12 20:18:25 +01:00
Michael Eischer	930a00ad54	checker: reuse bufio reader	2022-02-12 20:18:25 +01:00
Michael Eischer	f1e58e7c7f	checker: rewrite ReadData to stream packs	2022-02-12 20:18:25 +01:00
Michael Eischer	9aa2eff384	Add plumbing to calculate backend specific file hash for upload This enables the backends to request the calculation of a backend-specific hash. For the currently supported backends this will always be MD5. The hash calculation happens as early as possible, for pack files this is during assembly of the pack file. That way the hash would even capture corruptions of the temporary pack file on disk.	2021-08-04 22:17:46 +02:00
Alexander Neumann	04ca69cc78	Address issues reported by golint	2021-01-30 20:45:57 +01:00
Alexander Neumann	3c753c071c	errcheck: More error handling	2021-01-30 20:02:37 +01:00
Alexander Neumann	16313bfcc9	errcheck: Add error check for MergeFinalIndexes()	2021-01-30 20:02:37 +01:00
Alexander Neumann	bdfedf1f5b	Merge pull request #3173 from MichaelEischer/unify-index-loading Unify index loading	2021-01-28 13:50:42 +01:00
Michael Eischer	e2b0072441	check: add progress bar to the tree structure check	2021-01-28 11:10:50 +01:00
Michael Eischer	258ce0c1e5	parallel: report progress for StreamTrees This assigns an id to each tree root and then keeps track of how many tree loads (i.e. trees referenced for the first time) are pending per tree root. Once a tree root and its subtrees were fully processed there are no more pending tree loads and the tree root is reported as processed.	2021-01-28 11:08:43 +01:00
Michael Eischer	6e03f80ca2	check: Split the parallelized tree loader into a reusable component The actual code change is minimal	2021-01-28 11:08:43 +01:00
Michael Eischer	1d7bb01a6b	check: Cleanup tree loading and switch to use errgroup The helper methods are now wired up in the Structure method.	2021-01-28 11:08:43 +01:00
Alexander Weiss	2a1add7538	check: remove file size counter	2020-12-23 02:34:31 +01:00
Michael Eischer	96904f8972	check: extract parallel index loading	2020-12-22 22:36:18 +01:00
Alexander Weiss	26f85779be	Parallelize ForAllSnapshots	2020-12-06 05:09:58 +01:00
Alexander Weiss	aa7a5f19c2	Use BlobHandle in index methods	2020-11-22 20:41:12 +01:00
Alexander Weiss	a851c53cbe	Use PackSize in checker	2020-11-21 22:13:54 +01:00
Alexander Weiss	c3ddde9e7d	Return hdrSize in ListPack	2020-11-21 22:13:54 +01:00
Michael Eischer	1f43cac12d	check: Only track data blobs when unused blobs should be reported This improves the memory usage of check a lot as it now only has to track tree blobs when run using the default parameters.	2020-11-15 18:43:07 +01:00
Michael Eischer	6da66c15d8	check: Simplify referenced blob tracking The result is identical as long as the context in not canceled. However, in that case the result is incomplete anyways.	2020-11-15 18:42:55 +01:00
Michael Eischer	3500f9490c	check: Simplify blob status tracking UnusedBlobs now directly reads the list of existing blobs from the repository index. This removes the need for the blobStatusExists flag, which in turn allows converting the blobRefs map into a BlobSet.	2020-11-15 18:42:42 +01:00
Michael Eischer	b8c7543a55	check: Merge 'size could not be found' and 'not found in index' errors By construction these two errors always show up in pairs: 'size could not be found' is printed when the blob is not found in the repository index. That blob is also part of the `blobs` array. Later on, check iterates over that array and checks whether the blob is marked as existing. Which cannot be the case as that mark is generated by iterating over the repository index. The merged warning no longer reports the blob index within a file. That information could also be derived by printing the affected tree using `cat` and searching for the blob.	2020-11-15 18:41:50 +01:00
Alexander Weiss	17bb77b1f9	check: Also check blob length and offset	2020-11-14 00:42:49 +01:00
Alexander Weiss	80dcfca191	check: Check sizes computed from index and pack header	2020-11-14 00:42:49 +01:00
MichaelEischer	46d31ab86d	Merge pull request #3058 from greatroar/counter Replace restic.Progress with new progress.Counter (fixes two race conditions)	2020-11-09 22:19:09 +01:00

1 2

98 Commits