Skip to content

Add Opteryx - #135

Open
joocer wants to merge 1 commit into
ClickHouse:mainfrom
mabel-dev:add-opteryx
Open

joocer wants to merge 1 commit into
ClickHouse:mainfrom
mabel-dev:add-opteryx

Conversation

@joocer

@joocer joocer commented Oct 4, 2026

Copy link
Copy Markdown

What this adds

An entry for Opteryx, an embedded SQL query engine (pip install opteryx-core), at 1m, 10m and 100m.

Opteryx is stateless in this benchmark. There's no load step, no schema and no index. The five queries read the Bluesky NDJSON files directly with READ_JSONL('<dir>/*.jsonl'). The only preparation is decompressing the published .json.gz files, which is what the existing _files_json rows describe.

Results (m6i.8xlarge, Ubuntu 24.04, opteryx-core 0.9.153)

Size Hot (sum of 5 queries) Cold (sum of 5 queries) Documents read
1m 1.38 s 15.8 s 1,000,000
10m 2.50 s 183.0 s 9,999,994
100m 13.76 s 1,822.7 s 99,999,968

How it was run

  • main.sh follows the same structure as the other entries. install.sh installs Python 3.14 (deadsnakes) and pins opteryx-core==0.9.153 from PyPI.
  • prepare_data.sh decompresses the first N .json.gz files. total_size.sh reports the size of those decompressed files, which is exactly what is queried.
  • run_queries.sh: for each query, sync and drop_caches, then one Opteryx process runs the query 3 times. Run 1 is cold; runs 2 and 3 are hot. The timing covers the query only, not interpreter start-up. This matches how server-based entries keep a warm process between runs; no results are cached.
  • The disk was a 500 GiB gp3 volume at default settings (3,000 IOPS, 125 MB/s): the same baseline throughput as the larger gp3 volumes in the other entries. Cold runs are bound by that throughput.

Notes on the queries

  • ignore_errors => true skips records that are not valid JSON. A few Bluesky records are split by a raw newline or contain a raw control character, which is why 9,999,994 of 10m and 99,999,968 of 100m documents are read. These match the counts ClickHouse loads.
  • Q4 returns the timestamp (CAST(MIN(time_us) AS TIMESTAMP[us])). Q5 returns the span in milliseconds, truncated (CAST((MAX - MIN) / 1000 AS INTEGER)). This can differ by 1 ms from date_diff, which counts millisecond boundaries crossed.
  • At 1m, the results match ClickHouse's published query results for all five queries. Q1 prints records with no collection as None. The _query_results files are included for each size.

A suggestion

Opteryx has no load time, and its storage is the source files themselves. A stateless filter and load time metrics, like ClickBench's, would let readers compare entries that query files in place separately from entries that load data first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant