Science deserves a permanent memory.

Most scientific data outlives neither its grant nor its hard drive. We build archives that validate every deposit, version every record, and answer to people and machines. Open infrastructure, designed to outlast its builders.

To open an archive, request access.

Tell us your field and we'll reply by email. No spam.

The problem

Every field wants a PDB.
Almost none get one.

The PDB works because deposition is structured, validation is thorough, and every record is addressable. Building that infrastructure takes years and a consortium.

  1. 01

    Nobody can find it

    A DOI on a ZIP is not a queryable dataset. Reuse depends on guessing what the columns mean.

  2. 02

    Quality is assumed, not checked

    Without validation at deposition, missing metadata and broken ontology terms surface years later, in someone else's analysis.

  3. 03

    Unfit for training

    Training on public data means months of scraping and normalising, with no record of what was processed by which version of which script.

  4. 04

    The infrastructure outlives the grant

    Someone has to keep the server up after the postdoc leaves. Many portals degrade within three years.

Open source

We built OSA. We host it too.

Open Science Archive is an open source scientific data platform: deposition, validation, records, search and export. Amacrin hosts OSA archives, so you can run a data bank for your field without operating one.

Nothing about your archive is proprietary. Export everything at any time, or take the node in-house and run it yourself.

The platform

Write the model. Everything else follows.

Record schemas, validation checks, feature hooks. From that one file Amacrin derives the rest: deposition, builds, hosting, search and APIs.

  1. 01
    DepositThe deposition flow is generated from your schema. Depositors are their ORCID.
  2. 02
    ValidateEvery submission is checked before it exists. Failures stay private to the depositor.
  3. 03
    EnrichHooks attach embeddings, QC metrics, predictions. Versioned per build.
  4. 04
    MintEvery record gets a page, an SRN, and semantic search.
  5. 05
    ServeREST, bulk export, MCP. Any node can mirror or fork, provenance intact.

Structured Resource Names

Stable, versioned addresses that still resolve in ten years.

urn:osa:antibodies.example.org:ab:AB-04412@3

Your domain, your rules

Public reads by default. You decide who deposits and who curates.

antibodies.example.org

Nothing to operate

Compute, database, registry and TLS are ours to keep up. Export everything whenever you want.

tls · backups · upgrades
Agent-native

Built to be asked

Every archive speaks MCP and publishes a skill file describing what it holds. Ask in plain language; answers come back as charts and tables, not raw JSON.

semantic search · typed filters · exports pinned to a build

antibodies.example.org/mcp6 tools

Which VH germlines give the tightest PD-1 binders?

median affinity_kd_nM by ighv_germline · 4,120 of 7,204 records with paired germline call
11010010001,412IGHV3-233-23843IGHV1-691-691,107IGHV3-303-30402IGHV4-344-34356IGHV1-461-46ighv germlineaffinity kd nM (log)

IGHV3-23 has the lowest median KD, but 68% of those entries are bivalent IgG measured by SPR, where avidity inflates apparent affinity. Restricted to monovalent Fab, the ranking flattens and IGHV1-69 comes out ahead. Germline explains little here: CDR-H3 length and composition carry most of the signal.

Dataset cuts

Take the slice you need. Give it a permanent name.

Most work on an archive starts by narrowing it: the records that pass QC, the assays from one platform, everything since the last release. A cut is that subset, exported with its features computed and given a name that resolves to exactly these rows.

amacrin cut antibodies \  --schema antibody-sequence@2.1.0 \  --where 'target = "PD-1" and qc.status = "pass"' \  --features affinity_kd_nM, expression_yield, embedding \  --format parquet
idaffinity_kd_nMexpression_yieldembedding
AB-00417@20.42118float32[1280]
AB-01186@11.3796float32[1280]
AB-02940@30.09142float32[1280]
AB-04412@32.8571float32[1280]
AB-07356@10.61205float32[1280]
Rows
12,102
Format
parquet · 1.9 GB
Licence
CC-BY-4.0 · 41 labs
SRN
urn:osa:antibodies.example.org:cut:2f9a80@1

Filter on fields, not filenames

Query typed metadata. If a field is in the schema you can filter on it, and everything that matched passed the same checks.

Features arrive precomputed

Embeddings, QC metrics and predictions come out as columns, versioned with the build that produced them. Feed a model without rebuilding a pipeline.

One freeze for a whole consortium

Publish a cut as the official release and every member works from the same rows. Re-cut later and diff what changed.

Provenance

Every row remembers

Open any row and the record behind it answers: the specimen, the lab, the instrument, the licence, and everything that changed since.

cut 2f9a80urn:osa:antibodies.example.org:ab:AB-04412@3 opened
idaffinity_kd_nMexpression_yieldembedding
SpecimenIgG1 Fab, human framework VH3-23isolated from phage panel P-118 · library size 5×10⁹
AssaySurface plasmon resonanceBiacore 8K · triplicate injections · batch B-07
Deposited byExample Antibody Foundry0000-0000-0000-0000 · 04 Nov 2025
LicenceCC-BY-4.0redistribution permitted, attribution required
Versions@3 current · @2, @1 superseded@3 corrected CDR-H3 sequence, 12 Feb 2026
affinity_kd_nMDerived, not depositedcompute-binding-features · bld_8f3a91c2
SpecimenIgG1 Fab, human framework VH3-23isolated from phage panel P-118 · library size 5×10⁹
AssaySurface plasmon resonanceBiacore 8K · triplicate injections · batch B-07
Deposited byExample Antibody Foundry0000-0000-0000-0000 · 04 Nov 2025
LicenceCC-BY-4.0redistribution permitted, attribution required
Versions@3 current · @2, @1 superseded@3 corrected CDR-H3 sequence, 12 Feb 2026
affinity_kd_nMDerived, not depositedcompute-binding-features · bld_8f3a91c2
SpecimenIgG1 Fab, human framework VH3-23isolated from phage panel P-118 · library size 5×10⁹
AssaySurface plasmon resonanceBiacore 8K · triplicate injections · batch B-07
Deposited byExample Antibody Foundry0000-0000-0000-0000 · 04 Nov 2025
LicenceCC-BY-4.0redistribution permitted, attribution required
Versions@3 current · @2, @1 superseded@3 corrected CDR-H3 sequence, 12 Feb 2026
affinity_kd_nMDerived, not depositedcompute-binding-features · bld_8f3a91c2
SpecimenIgG1 Fab, human framework VH3-23isolated from phage panel P-118 · library size 5×10⁹
AssaySurface plasmon resonanceBiacore 8K · triplicate injections · batch B-07
Deposited byExample Antibody Foundry0000-0000-0000-0000 · 04 Nov 2025
LicenceCC-BY-4.0redistribution permitted, attribution required
Versions@3 current · @2, @1 superseded@3 corrected CDR-H3 sequence, 12 Feb 2026
affinity_kd_nMDerived, not depositedcompute-binding-features · bld_8f3a91c2
SpecimenIgG1 Fab, human framework VH3-23isolated from phage panel P-118 · library size 5×10⁹
AssaySurface plasmon resonanceBiacore 8K · triplicate injections · batch B-07
Deposited byExample Antibody Foundry0000-0000-0000-0000 · 04 Nov 2025
LicenceCC-BY-4.0redistribution permitted, attribution required
Versions@3 current · @2, @1 superseded@3 corrected CDR-H3 sequence, 12 Feb 2026
affinity_kd_nMDerived, not depositedcompute-binding-features · bld_8f3a91c2
NOTEThe same lookup works from the API and from an agent, so a model can check what it's looking at.
  • Records are versioned, never replacedA correction publishes a new version and keeps the old one addressable. An analysis that used @1 still resolves to exactly what it saw.
  • Corrections find your cutsWhen a depositor revises a record, every cut that contained it is flagged. You find out which models were trained on the old value.
  • Licence terms stay attachedWhat you are allowed to do with a record travels with it through export, mirroring and forking.
Audiences

Who runs archives

Start with one dataset, or bring the whole field

We're onboarding archives one field at a time. Leave your email and what you'd deposit, and we'll work with you to plan and set up your archive.

Tell us your field and we'll reply by email to help plan your archive. No spam.

Running a consortium or a large release? Talk to us directly.