Repository navigation
Add Opteryx - #135
Open
joocer wants to merge 1 commit into
Open
Add Opteryx#135joocer wants to merge 1 commit into
joocer wants to merge 1 commit into
Conversation
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this adds
An entry for Opteryx, an embedded SQL query engine (
pip install opteryx-core), at 1m, 10m and 100m.Opteryx is stateless in this benchmark. There's no load step, no schema and no index. The five queries read the Bluesky NDJSON files directly with
READ_JSONL('<dir>/*.jsonl'). The only preparation is decompressing the published.json.gzfiles, which is what the existing_files_jsonrows describe.Results (m6i.8xlarge, Ubuntu 24.04, opteryx-core 0.9.153)
How it was run
main.shfollows the same structure as the other entries.install.shinstalls Python 3.14 (deadsnakes) and pinsopteryx-core==0.9.153from PyPI.prepare_data.shdecompresses the first N.json.gzfiles.total_size.shreports the size of those decompressed files, which is exactly what is queried.run_queries.sh: for each query,syncanddrop_caches, then one Opteryx process runs the query 3 times. Run 1 is cold; runs 2 and 3 are hot. The timing covers the query only, not interpreter start-up. This matches how server-based entries keep a warm process between runs; no results are cached.Notes on the queries
ignore_errors => trueskips records that are not valid JSON. A few Bluesky records are split by a raw newline or contain a raw control character, which is why 9,999,994 of 10m and 99,999,968 of 100m documents are read. These match the counts ClickHouse loads.CAST(MIN(time_us) AS TIMESTAMP[us])). Q5 returns the span in milliseconds, truncated (CAST((MAX - MIN) / 1000 AS INTEGER)). This can differ by 1 ms fromdate_diff, which counts millisecond boundaries crossed.None. The_query_resultsfiles are included for each size.A suggestion
Opteryx has no load time, and its storage is the source files themselves. A
statelessfilter and load time metrics, like ClickBench's, would let readers compare entries that query files in place separately from entries that load data first.