> ## Documentation Index
> Fetch the complete documentation index at: https://instance.bio/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Quickstart

> Load, inspect, rank, and validate Instance Parquet datasets using Python and SQL.

Instance datasets are delivered as Parquet files organized into table-specific folders (for example, `./data/binding_scores/binding_scores_20260103.parquet`). You can load a table by passing its directory path.

## Environment Setup

To set up an isolated Python virtual environment using [`uv`](https://docs.astral.sh/uv/):

```bash theme={null}
# Create and activate an isolated virtual environment
uv venv
source .venv/bin/activate

# Install your preferred data analysis library
uv pip install pandas polars duckdb pyarrow
```

## Read Dataset Tables

Store the path to your downloaded dataset in a `data_dir` variable so your scripts remain portable across environments. Pointing the reader to a table folder with a trailing slash loads the Parquet files in that folder. The examples below load the table and display the first 5 records.

<Tabs>
  <Tab title="Pandas">
    ```python theme={null}
    import pandas as pd

    data_dir = "./data"

    # Read the binding_scores folder directly and display the first 5 records
    scores = pd.read_parquet(f"{data_dir}/binding_scores/")
    print(scores.head(5))
    ```
  </Tab>

  <Tab title="Polars">
    ```python theme={null}
    import polars as pl

    data_dir = "./data"

    # Read the binding_scores folder and display the first 5 records
    scores = pl.read_parquet(f"{data_dir}/binding_scores/")
    print(scores.head(5))
    ```
  </Tab>

  <Tab title="DuckDB">
    ```python theme={null}
    import duckdb

    data_dir = "./data"

    # Query the Parquet folder directly with SQL and display the first 5 records
    scores = duckdb.sql(f"select * from read_parquet('{data_dir}/binding_scores/') limit 5").df()
    print(scores)
    ```
  </Tab>
</Tabs>

## Screen and Rank Candidates by Binding Score

To prioritize binders, filter `binding_scores` for a designated target antigen, select candidates that were expressed (`is_expressed = true`) and designed for the target (`targeting = true`), and rank the candidates in descending order by binding score. The examples below rank candidates and display the first 10 top-ranking records.

<Tabs>
  <Tab title="Pandas">
    ```python theme={null}
    import pandas as pd

    data_dir = "./data"
    target_antigen_gene_name = "YOUR_ANTIGEN_GENE_NAME"  # Matches antigen_gene_name

    scores = pd.read_parquet(f"{data_dir}/binding_scores/")

    ranked_candidates = (
        scores.loc[
            (scores["antigen_gene_name"] == target_antigen_gene_name)
            & scores["is_expressed"]
            & scores["targeting"]
        ]
        .sort_values(by="binding_score_50nM", ascending=False)
    ).head(10)

    # Display the first 10 top-ranking candidates
    print(ranked_candidates)
    ```
  </Tab>

  <Tab title="Polars">
    ```python theme={null}
    import polars as pl

    data_dir = "./data"
    target_antigen_gene_name = "YOUR_ANTIGEN_GENE_NAME"

    ranked_candidates = (
        pl.read_parquet(f"{data_dir}/binding_scores/")
        .filter(
            (pl.col("antigen_gene_name") == target_antigen_gene_name)
            & pl.col("is_expressed")
            & pl.col("targeting")
        )
        .sort(by="binding_score_50nM", descending=True, nulls_last=True)
    ).head(10)

    # Display the first 10 top-ranking candidates
    print(ranked_candidates)
    ```
  </Tab>

  <Tab title="DuckDB">
    ```python theme={null}
    import duckdb

    data_dir = "./data"
    target_antigen_gene_name = "YOUR_ANTIGEN_GENE_NAME"

    # Query and display the first 10 top-ranking candidates
    ranked_candidates = duckdb.execute(f"""
    select *
    from read_parquet('{data_dir}/binding_scores/')
    where antigen_gene_name = ?
      and is_expressed = true
      and targeting = true
    order by binding_score_50nM desc
    limit 10
    """, [target_antigen_gene_name]).df()

    print(ranked_candidates)
    ```
  </Tab>
</Tabs>

<Note>
  **Comparing Computational Model Predictions:** If your team generated predictions before
  submitting your order, you can join your predictions table with `binding_scores` using
  `candidate_name` to compute rank correlation (such as Spearman's $\rho$) and evaluate model
  performance against experimental ground truth.
</Note>

## Inspect Replicate Measurements and UMI Counts

If you want to inspect replicate scores for your top rankers, you can do so using `replicate_binding_scores` to see per-sample measurements and molecule counts (`binding_umi_count`). The examples below filter for replicate measurements corresponding to the top candidates identified above.

<Tabs>
  <Tab title="Pandas">
    ```python theme={null}
    import pandas as pd

    data_dir = "./data"

    # Reuse the top candidates from ranked_candidates above
    key_cols = ["candidate_library_id", "sino_catalog_id", "variant_id"]
    top_candidates = ranked_candidates[key_cols]

    replicate_binding_scores = pd.read_parquet(f"{data_dir}/replicate_binding_scores/")

    lead_replicates = (
        replicate_binding_scores[
            replicate_binding_scores.set_index(key_cols).index.isin(
                pd.MultiIndex.from_frame(top_candidates)
            )
        ]
        .sort_values(by=["candidate_id", "antigen_concentration_nM"])
    )

    print(lead_replicates)
    ```
  </Tab>

  <Tab title="Polars">
    ```python theme={null}
    import polars as pl

    data_dir = "./data"

    replicate_binding_scores = pl.scan_parquet(f"{data_dir}/replicate_binding_scores/")

    lead_replicates = (
        replicate_binding_scores
        .join(
            ranked_candidates.lazy(),
            on=["candidate_library_id", "sino_catalog_id", "variant_id"],
            how="semi",
        )
        .sort(by=["candidate_id", "antigen_concentration_nM"])
        .collect()
    )

    print(lead_replicates)
    ```
  </Tab>

  <Tab title="DuckDB">
    ```python theme={null}
    import duckdb

    data_dir = "./data"

    lead_replicates = duckdb.execute(f"""
        select r.*
        from read_parquet('{data_dir}/replicate_binding_scores/') as r
        where exists (
            select 1
            from ranked_candidates as c
            where c.candidate_library_id = r.candidate_library_id
            and c.sino_catalog_id = r.sino_catalog_id
            and c.variant_id = r.variant_id
        )
        order by r.candidate_id, r.antigen_concentration_nM
    """).df()

    print(lead_replicates)
    ```
  </Tab>
</Tabs>

## Best Practices for Joining Tables

When joining Instance tables, following standard join hygiene ensures data integrity and prevents unintended duplicate rows.

### 1. Join on a Table's Complete Unique Grain

Always join two tables on the complete unique grain of at least one of the tables. Joining only on a subset of identifying keys (such as joining on `candidate_id` alone when multiple variants or concentrations exist) can lead to unintentional row duplication.

| Table | Unique Grain Columns |
| :- | :- |
| `binding_scores` | `(candidate_library_id, sino_catalog_id, variant_id)` |
| `specificity_scores` | `(candidate_library_id, sino_catalog_id, variant_id)` |
| `replicate_binding_scores` | `(expression_sample_id, binding_sample_id, variant_id)` |
| `replicate_specificity_scores` | `(expression_sample_id, binding_sample_id, variant_id)` |
| `dim_samples` | `sample_id` |
| `dim_candidates` | `(sample_id, sino_catalog_id, antigen_concentration_nM)` |
| `aligned_reads` | `(read_id, sample_id)` |
| `binding_umi_counts` | `(binding_sample_id, variant_id, umi)` |

### 2. Avoid Unnecessary Joins

Our score datasets (`binding_scores` and `specificity_scores`) are intentionally denormalized to include candidate names, full amino acid sequences, antigen catalog IDs, and antigen gene names. You do not need to join `dim_samples` or `dim_candidates` merely to access candidate sequences or target antigen names. Reserve table joins for cross-assay analysis, such as comparing `binding_scores` against `specificity_scores`.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.