Search before asking
Motivation
Repeated scans of the same snapshot and bucket still read and decode the manifest lists before checking the snapshot live-manifest-entry cache. The cached live entries already contain the state needed for that bucket, so this work adds avoidable CPU and metadata access to repeated point-read planning.
To reproduce, use the public orc/append_09.db/append_09 test fixture with snapshot 5 and bucket 0, enable the snapshot live-manifest-entry cache, and reuse one cache across two scans. Instrument the MANIFEST cache accesses, as in AppendOnlyFileStoreScanTest.TestSnapshotLiveManifestCachePath. Before the proposed change, the second scan revisits the manifest-list path even though its exact snapshot/bucket live-entry lookup hits.
Expected: an eligible exact hit bypasses manifest-list reads and decoding, while each request still applies its own filters. A miss must rebuild the complete bucket. Snapshot selection and the existing snapshot-identity validation must remain intact.
Solution
Resolve the snapshot first and consult the existing snapshot live-entry cache before reading manifest lists. Retain the table file count in the transient cached payload so skipped-file metrics remain correct; report zero scanned manifests for an exact hit. Version the transient payload so incompatible entries fall back to rebuilding.
Draft implementation: #374. Regression tests check the manifest-cache access count, changed predicates, eviction/rebuild, and retained file counts. This change does not introduce a filesystem block cache or alter the persisted table format.
Anything else?
The identical patch passed 88 focused tests in a downstream C++17 Release build. Upstream CI for the draft still requires maintainer approval; no standalone end-to-end performance uplift is claimed for this individual optimization.
Are you willing to submit a PR?
Search before asking
Motivation
Repeated scans of the same snapshot and bucket still read and decode the manifest lists before checking the snapshot live-manifest-entry cache. The cached live entries already contain the state needed for that bucket, so this work adds avoidable CPU and metadata access to repeated point-read planning.
To reproduce, use the public
orc/append_09.db/append_09test fixture with snapshot 5 and bucket 0, enable the snapshot live-manifest-entry cache, and reuse one cache across two scans. Instrument theMANIFESTcache accesses, as inAppendOnlyFileStoreScanTest.TestSnapshotLiveManifestCachePath. Before the proposed change, the second scan revisits the manifest-list path even though its exact snapshot/bucket live-entry lookup hits.Expected: an eligible exact hit bypasses manifest-list reads and decoding, while each request still applies its own filters. A miss must rebuild the complete bucket. Snapshot selection and the existing snapshot-identity validation must remain intact.
Solution
Resolve the snapshot first and consult the existing snapshot live-entry cache before reading manifest lists. Retain the table file count in the transient cached payload so skipped-file metrics remain correct; report zero scanned manifests for an exact hit. Version the transient payload so incompatible entries fall back to rebuilding.
Draft implementation: #374. Regression tests check the manifest-cache access count, changed predicates, eviction/rebuild, and retained file counts. This change does not introduce a filesystem block cache or alter the persisted table format.
Anything else?
The identical patch passed 88 focused tests in a downstream C++17 Release build. Upstream CI for the draft still requires maintainer approval; no standalone end-to-end performance uplift is claimed for this individual optimization.
Are you willing to submit a PR?