Skip to main content
Instance datasets are delivered as Parquet files. Every dataset provides a structured schema designed for downstream machine learning workflows, statistical evaluation, and assay validation. For code examples on loading, summarizing, and ranking datasets using Python (Pandas, Polars) and SQL (DuckDB), see the Quickstart guide.

Assay and Data Lifecycle

Every dataset is the outcome of a five-stage workflow that transforms customer sequence designs into quantitative biological measurements.
  1. Order: You submit your target antigens and candidate design sequences that you want labels for.
  2. Planning: The assay layout is configured—defining which libraries, antigens, concentrations, and other assay conditions will be tested. Identifiers containing _plan_ (such as candidate_library_plan_id and sample_plan_id) are created during this stage to track the planned blueprint.
  3. Execution: The assay is run in the wet lab under the planned conditions, through candidate expression, antigen binding selection, and sequencing. Concrete identifiers (candidate_library_id and sample_id) are assigned to the physical reaction vessels.
  4. Dataset Preparation: Sequencing data is processed: reads are aligned to candidate sequences, PCR duplicates are removed using Unique Molecular Identifiers (UMIs), and final scores and counts are formatted into Parquet tables.

Process Overview

The assay workflow measures candidate abundance at sequential stages of the process. Material is sequenced at each stage, and score datasets quantify the relative abundance of candidates between stages.
Assay and Score Process Overview
  1. Pre-Expression (🧬 Manufacture Pools): Candidate sequence designs are manufactured into pooled DNA libraries (candidate_library_id). Material is sequenced here to establish baseline input DNA abundance.
  2. Post-Expression (🪢 Protein Expression): DNA pools are transcribed and translated in vitro to produce folded protein or peptide candidates. Material is sequenced as expression samples (expression_sample_id, selection = false). Comparing candidate abundance in Post-Expression relative to Pre-Expression yields the expression_scores.
  3. Post-Binding (🧪 Antigen Treatment): The expressed protein pool serves as the pre-binding baseline and is incubated with target antigens across tested concentrations. Material that successfully binds is captured and sequenced as binding samples (binding_sample_id, selection = true). Comparing candidate abundance in Post-Binding relative to Post-Expression yields the binding_scores (and sample-level replicate_binding_scores).
  4. Target vs. Off-Target Specificity (🎯 Specificity Scores): Specificity is evaluated by measuring candidate binding across multiple antigens—both intended on-target antigens and off-target counter-targets. Comparing a candidate’s binding on its intended target against its binding on off-target antigens directly from the post-binding data yields the specificity_scores (and sample-level replicate_specificity_scores).

Score Datasets (Final Analysis Stage)

Depending on the service and assay goals, Instance produces primary score datasets that summarize candidate performance. These tables represent the final modeled data products used for candidate ranking, hit discovery, and machine learning model training.

binding_scores

Quantitative binding scores measuring candidate interaction with target antigens across tested concentrations, delivered as summarized averages and replicate measurements.

specificity_scores

Selectivity measurements comparing candidate binding between target antigens and counter-target antigens across tested concentrations.

Supporting Datasets (Intermediate & Assay-Level Inspection)

To inspect data one step prior to the final statistical score summaries, Instance provides supporting datasets. These tables allow customers to evaluate physical sample metadata, trace submitted candidate sequences to assay conditions, inspect individual aligned sequencing reads, and audit umi-level counts directly without model aggregation.

dim_samples

Master reference table defining physical sample metadata, candidate libraries, antigens, concentrations, translation conditions, and selection status.

dim_candidates

Mapping table connecting customer-submitted candidate designs, names, and sequences to assay runs and antigens.

aligned_reads

Sequencing reads aligned to candidate sequences alongside their extracted Unique Molecular Identifiers (UMIs).

binding_umi_counts

UMI-level sequencing read counts for candidates and observed sequence variants across samples that underwent antigen selection.

Glossary

  • Unique Molecular Identifier (UMI): A short, randomized nucleotide sequence attached to each individual cDNA molecule before PCR amplification. During sequencing preparation, PCR generates many identical copies of each original molecule. By grouping reads that share the same UMI, PCR amplification duplicates are deduplicated so that each row represents a single physical molecule, preventing amplification bias from skewing count data.
  • Sequencer: The instrument that reads the nucleotide bases of amplified DNA constructs, measures base-calling quality, and outputs sequencing reads with unique instrument-assigned read identifiers (read_id).
  • Candidate Design (candidate_id): A content-addressed digest of the customer-submitted, intended amino acid or nucleotide sequence. Two designs with identical sequences share the exact same candidate_id across all datasets.
  • Observed Sequence Variant (variant_id): A content-addressed digest of the sequence variant actually detected after deep sequencing. Biological synthesis and replication can introduce minor sequence alterations. When the sequenced molecule matches the customer design perfectly, variant_id equals candidate_id.
  • Sample (sample_id): A distinct physical reaction vessel (such as a tube or well) containing an expressed candidate library exposed to a defined antigen concentration and biochemical condition. Uniquely identified by sample_id.
  • Planned Identifiers vs. Executed Identifiers: Identifiers containing _plan_ (such as sample_plan_id and candidate_library_plan_id) denote virtual specifications established during the Planning stage. Identifiers without _plan_ (sample_id and candidate_library_id) denote physical reagents, reaction wells, and executed runs in the Execution stage.