> ## Documentation Index
> Fetch the complete documentation index at: https://instance.bio/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# aligned_reads

Contains individual sequencing reads that aligned a candidate sequence in the library. The Unique Molecular Identifier (UMI) is parsed from each read. During the course of our assay, we use PCR to amplify the amount of bound material before sequencing. PCR amplification produces multiple identical copies of each original molecule, which copies the UMIs at different rates leading to biases in observed read counts. By deduplicating the UMIs observed in the reads, the initial physical molecule can be distinguished from independently captured molecules, removing the amplification bias.

**Grain:** one row per `(read_id, sample_id)`.

| Column | Type | Description |
| :- | :- | :- |
| `read_id` | `string` | Identifier assigned to the sequencing read by the sequencing instrument. |
| `sample_id` | `string` | Identifies the physical sample the read came from. |
| `candidate_id` | `string` | Identifies the candidate this read aligned to across the dataset. Two candidates with identical sequences share the same `candidate_id`. |
| `read_sequence` | `string` | The observed sequence of the read as recorded by the sequencer. |
| `variant_id` | `string` | Identifies the specific sequence variant observed for this read. Equals `candidate_id` when the observed sequence matches the candidate sequence exactly. |
| `umi` | `string` | Unique Molecular Identifier (UMI) sequence extracted from the read, identifying the single original physical molecule before PCR amplification. |


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.