About this project
Opteryx Core is the SQL execution engine behind opteryx.app, published as a fork of Opteryx with a smaller, more opinionated API and configuration surface shaped around the workloads of the hosted service. It is designed for fast, read-heavy analytical queries over columnar data: it handles SQL parsing, planning, predicate pushdown, projection pruning and execution, so datasets can be queried from Python without standing up a separate warehouse.
Architecture
Query planning is written in Python; query execution is native. Once the planner produces a physical plan, the engine runs it in compiled code end to end — scan, operators, scheduling and dispatch — and neither PyArrow nor NumPy is present anywhere in the engine. Results are returned as Draken morsels, batches of columns streamed as the engine produces them, so a large result does not have to fit in memory at once. A morsel exposes num_rows, column_names and column(name).to_pylist(); after the stream is read to the end, session.rowcount reports the number of rows delivered.
Getting started
Requirements are Python 3.11 or later, a C/C++ toolchain for local source builds, and Rust/Cargo for the Rust extension. Install with pip install opteryx-core and import it as opteryx. A minimal local example registers a workspace with DiskConnector and queries dot-separated dataset names resolved relative to the current working directory, for example data.planets resolving to ./data/planets, with the format detected from file extensions. A command line is also available via python -m opteryx for querying without writing Python.
Intended uses
The project lists powering the execution layer used by opteryx.app, running analytical SQL against local Parquet, CSV, JSONL and .skene datasets, embedding a query engine inside Python applications, scripts, notebooks and services, working on engine internals such as planning, native execution and file-format performance, and using the file engine or the .skene format on their own through the rugo and libskene wheels.
Repository layout and distributions
The repository contains the SQL engine (opteryx/), the native columnar vector substrate and morsels (draken/), the file engine for Parquet, CSV and JSONL read and write (rugo/), the .skene columnar file format with C++ reader, writer and normative specification (skene/), compute extension sources in Rust and C++ (src/), generated catalog snapshots (reference/), tests, testdata, docs, development scripts, vendored dependencies and shared build machinery in build_common.py.
One source tree produces three wheels, single-sourced in build_common.py so they cannot drift: opteryx-core (imported as opteryx) bundles the full SQL engine with draken, rugo and skene; rugo provides the file engine plus draken for reading and writing files without the SQL engine; libskene (imported as skene) provides the .skene reader and writer plus draken. draken is not published separately. rugo and skene are parallel and neither depends on the other. Wheels are built in CI, not locally; local development uses the Makefile targets such as make dev-install, make compile, make c, make q, make test, make dt and make check.
File formats
Datasets are read by extension and a dataset is one format throughout; a directory mixing formats is an error rather than a best-effort read. Parquet is the default for stored data and interchange, read through rugo, as are CSV and JSONL/NDJSON. The .skene format is draken-native: it stores one or more row groups of draken vectors losslessly, including refinements Parquet drops — an IPv4 column round-trips as a UINT32 refined by an IPV4 logical descriptor, and dictionary encoding and layout hints are restored rather than re-derived. It is deliberately not portable and no foreign reader is promised, so Parquet remains the choice for interchange. Parquet, CSV and JSONL files can also be named directly with the read_parquet(), read_csv() and read_jsonl() table functions; there is no read_skene().
Catalog integration
Opteryx Core is described as working best when paired with the opteryx_catalog library, which is the intended model for named datasets, catalog-backed tables and the general experience used in opteryx.app. Configuration sets a default connector with catalog, Firestore project and database, and a GCS bucket, after which catalog-backed datasets can be queried with dot-separated names such as public.space.planets. For local data, registered workspaces such as testdata, scratch or data are typical.
Positioning and contributing
The project presents itself as an embedded analytical engine rather than a full end-user platform: for a hosted experience and multi-tenant service features, opteryx.app is the recommended route, while this package provides the core engine directly. Contributions are invited in the form of usage on personal datasets, bug reports when queries, schemas or performance misbehave, pull requests for fixes, tests, docs or performance, and shared repro cases, failing queries and edge-case Parquet files. The project is licensed under Apache-2.0, with documentation at docs.opteryx.app.
Comments
0 Rating appears after 10 ratings
Sign in to join the discussion.