Assay and Data Lifecycle
Every dataset is the outcome of a five-stage workflow that transforms customer sequence designs into quantitative biological measurements.- Order: You submit your target antigens and candidate design sequences that you want labels for.
- Planning: The assay layout is configured—defining which libraries, antigens, concentrations, and other assay conditions will be tested. Identifiers containing
_plan_(such ascandidate_library_plan_idandsample_plan_id) are created during this stage to track the planned blueprint. - Execution: The assay is run in the wet lab under the planned conditions, through candidate expression, antigen binding selection, and sequencing. Concrete identifiers (
candidate_library_idandsample_id) are assigned to the physical reaction vessels. - Dataset Preparation: Sequencing data is processed: reads are aligned to candidate sequences, PCR duplicates are removed using Unique Molecular Identifiers (UMIs), and final scores and counts are formatted into Parquet tables.
Process Overview
The assay workflow measures candidate abundance at sequential stages of the process. Material is sequenced at each stage, and score datasets quantify the relative abundance of candidates between stages.- Pre-Expression (🧬 Manufacture Pools): Candidate sequence designs are manufactured into pooled DNA libraries (
candidate_library_id). Material is sequenced here to establish baseline input DNA abundance. - Post-Expression (🪢 Protein Expression): DNA pools are transcribed and translated in vitro to produce folded protein or peptide candidates. Material is sequenced as expression samples (
expression_sample_id,selection = false). Comparing candidate abundance in Post-Expression relative to Pre-Expression yields theexpression_scores. - Post-Binding (🧪 Antigen Treatment): The expressed protein pool serves as the pre-binding baseline and is incubated with target antigens across tested concentrations. Material that successfully binds is captured and sequenced as binding samples (
binding_sample_id,selection = true). Comparing candidate abundance in Post-Binding relative to Post-Expression yields thebinding_scores(and sample-levelreplicate_binding_scores). - Target vs. Off-Target Specificity (🎯 Specificity Scores): Specificity is evaluated by measuring candidate binding across multiple antigens—both intended on-target antigens and off-target counter-targets. Comparing a candidate’s binding on its intended target against its binding on off-target antigens directly from the post-binding data yields the
specificity_scores(and sample-levelreplicate_specificity_scores).
Score Datasets (Final Analysis Stage)
Depending on the service and assay goals, Instance produces primary score datasets that summarize candidate performance. These tables represent the final modeled data products used for candidate ranking, hit discovery, and machine learning model training.binding_scores
Quantitative binding scores measuring candidate interaction with target antigens across
tested concentrations, delivered as summarized averages and replicate measurements.
specificity_scores
Selectivity measurements comparing candidate binding between target antigens and
counter-target antigens across tested concentrations.
Supporting Datasets (Intermediate & Assay-Level Inspection)
To inspect data one step prior to the final statistical score summaries, Instance provides supporting datasets. These tables allow customers to evaluate physical sample metadata, trace submitted candidate sequences to assay conditions, inspect individual aligned sequencing reads, and audit umi-level counts directly without model aggregation.dim_samples
Master reference table defining physical sample metadata, candidate libraries, antigens,
concentrations, translation conditions, and selection status.
dim_candidates
Mapping table connecting customer-submitted candidate designs, names, and sequences to assay
runs and antigens.
aligned_reads
Sequencing reads aligned to candidate sequences alongside their extracted Unique Molecular
Identifiers (UMIs).
binding_umi_counts
UMI-level sequencing read counts for candidates and observed sequence variants across
samples that underwent antigen selection.
Glossary
- Unique Molecular Identifier (UMI): A short, randomized nucleotide sequence attached to each individual cDNA molecule before PCR amplification. During sequencing preparation, PCR generates many identical copies of each original molecule. By grouping reads that share the same UMI, PCR amplification duplicates are deduplicated so that each row represents a single physical molecule, preventing amplification bias from skewing count data.
- Sequencer: The instrument that reads the nucleotide bases of amplified DNA constructs, measures base-calling quality, and outputs sequencing reads with unique instrument-assigned read identifiers (
read_id). - Candidate Design (
candidate_id): A content-addressed digest of the customer-submitted, intended amino acid or nucleotide sequence. Two designs with identical sequences share the exact samecandidate_idacross all datasets. - Observed Sequence Variant (
variant_id): A content-addressed digest of the sequence variant actually detected after deep sequencing. Biological synthesis and replication can introduce minor sequence alterations. When the sequenced molecule matches the customer design perfectly,variant_idequalscandidate_id. - Sample (
sample_id): A distinct physical reaction vessel (such as a tube or well) containing an expressed candidate library exposed to a defined antigen concentration and biochemical condition. Uniquely identified bysample_id. - Planned Identifiers vs. Executed Identifiers: Identifiers containing
_plan_(such assample_plan_idandcandidate_library_plan_id) denote virtual specifications established during the Planning stage. Identifiers without_plan_(sample_idandcandidate_library_id) denote physical reagents, reaction wells, and executed runs in the Execution stage.

