Skip to content

Add Ranking for Queries Without Defined Subpath Types or Indexes #90

Description

@eldenmoon

Hello guys, I believe that semi-structured data is highly flexible and dynamic, and its schema and types are often difficult to anticipate before table creation (which also makes it hard to create indexes in advance). Therefore, shouldn’t we provide a ranking — ideally the main ranking — that shows performance without defining any subpath types or subpath indexes?

Activity

  1. rschu1ze commented on Sep 16, 2025

    @rschu1ze
    Member

    The blog post describes the design choices in JSONBench.

    I can only speak for ClickHouse here. By "subpath types" or "subpath indexes", I guess you mean in the DDL file

    CREATE TABLE bluesky
    (
        `data` JSON(
            kind LowCardinality(String),
            commit.operation LowCardinality(String),
            commit.collection LowCardinality(String),
            did String,
            time_us UInt64)  CODEC(ZSTD(1))
    )
    ORDER BY (
        data.kind,
        data.commit.operation,
        data.commit.collection,
        data.did,
        fromUnixTimestamp64Micro(data.time_us));

    the parameters of the JSON type, right? I agree that this is certain level of tuning (paths kind, commit.operation, commit.collection, did and time_us will be stored as separate columns). One can argue either way if that is okay or not (e.g. with rather static dashboard queries, this amount of tuning may be fine). Anyways, I guess it makes sense to introduce a configuration without any tuning as well.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions