Open data platform for traceable, multimodal AI
Improve your workflows with open-source context, memory & data access.
Query, trace & govern with a lineage-native lakehouse that works across files, arrays, ontologies & notes to make agents more efficient and trustworthy.
Manage biological data through modules for bio-formats & registries — by the
creators of



Lineage
Trace data, code & agents
Track where data comes from and what it's used for, no matter how you run your code.
Lakehouse beyond tables
Query across many datasets
Query and batch-load datasets with lakehouse support for a wide range of table & array formats. Manage their features & schemas as metadata in Postgres or SQLite.


Notes, LIMS, ELN
Unify records management
Manage metadata in notes, registries & relational sheets in sync with datasets in storage. Use a single Python/R class with built-in ontologies.

FAIR datasets
Validate & annotate datasets
Use schemas to enforce consistency across your datasets. Annotate with a single line of code.
Branching & versioning
Manage changes
Manage changes to datasets, records, and models like you manage changes to software with git. Merge Change Requests from agents and collaborators. Co-version data and code.
Zero lock-in
Administer with ease while staying in control
Manage fine-grained permissions for humans & agents with SaaS-like simplicity directly at the database and storage level. Do not give up admin control on AWS, GCP, or in your own infrastructure.

Context
Build your organization's long-term memory
As team & agents work, data, models & reports get mapped into the lakehouse — building recursively queryable memory & training data that compounds over time.