Skip to main content
Instance datasets are delivered as Parquet files organized into table-specific folders (for example, ./data/binding_scores/binding_scores_20260103.parquet). You can load a table by passing its directory path.

Environment Setup

To set up an isolated Python virtual environment using uv:

Read Dataset Tables

Store the path to your downloaded dataset in a data_dir variable so your scripts remain portable across environments. Pointing the reader to a table folder with a trailing slash loads the Parquet files in that folder. The examples below load the table and display the first 5 records.

Screen and Rank Candidates by Binding Score

To prioritize binders, filter binding_scores for a designated target antigen, select candidates that were expressed (is_expressed = true) and designed for the target (targeting = true), and rank the candidates in descending order by binding score. The examples below rank candidates and display the first 10 top-ranking records.
Comparing Computational Model Predictions: If your team generated predictions before submitting your order, you can join your predictions table with binding_scores using candidate_name to compute rank correlation (such as Spearman’s ρ\rho) and evaluate model performance against experimental ground truth.

Inspect Replicate Measurements and UMI Counts

If you want to inspect replicate scores for your top rankers, you can do so using replicate_binding_scores to see per-sample measurements and molecule counts (binding_umi_count). The examples below filter for replicate measurements corresponding to the top candidates identified above.

Best Practices for Joining Tables

When joining Instance tables, following standard join hygiene ensures data integrity and prevents unintended duplicate rows.

1. Join on a Table’s Complete Unique Grain

Always join two tables on the complete unique grain of at least one of the tables. Joining only on a subset of identifying keys (such as joining on candidate_id alone when multiple variants or concentrations exist) can lead to unintentional row duplication.

2. Avoid Unnecessary Joins

Our score datasets (binding_scores and specificity_scores) are intentionally denormalized to include candidate names, full amino acid sequences, antigen catalog IDs, and antigen gene names. You do not need to join dim_samples or dim_candidates merely to access candidate sequences or target antigen names. Reserve table joins for cross-assay analysis, such as comparing binding_scores against specificity_scores.