Protein Plant Synthesis

If you had to produce a therapeutic protein in a plant, which host and which DNA design should you try first?

A leakage-controlled surrogate ML pipeline that designs DNA for therapeutic proteins across four plant hosts, and explains its own predictions.

4 hostsArabidopsis, rice, maize and date palm, compared under protein-grouped evaluation
Year
2025–2026
Status
Working pipeline + dashboard · simulated target, not wet-lab data · 26 tests
Stack
Python, scikit-learn, XGBoost, LightGBM, SHAP, pandas, Streamlit, Plotly, py3Dmol
01

From UniProt to a ranked decision

Proteins are fetched from UniProt and quality-checked. Each one gets a naive back-translation and host-aware optimised DNA for four plant hosts, using host-biased codon tables. Structure descriptors come from AlphaFold, with a deterministic fallback so the app still runs offline.

Codon adaptation (CAI), GC compatibility, rare-codon ratios and protein features feed a set of regressors. The final screen turns predictions into an operations view: which host to try, at what estimated cost and time.

System architecture: data collection, DNA generation and feature engineering, proxy target construction
02

Making the evaluation credible

Grouped splits

The same protein appears once per host. Random splits would leak an accession into both train and test, so splits are grouped by protein (GroupShuffleSplit, GroupKFold).

Beat the trivial model first

Every model is compared against a dummy baseline that predicts the mean. The dashboard asks the question explicitly: is this better than a trivial prediction?

Explain, then trust

Permutation importance, local SHAP explanations and residual analysis are part of the product, not an afterthought.

One more section for engineers: architecture, hyper-parameters and design decisions.

03

Under the hood

Under the hood
Models
Dummy baseline, Linear, Ridge, Random Forest (350 trees, depth 14), XGBoost (350, depth 6, lr 0.05), optional LightGBM
Features
CAI, GC content, codon diversity, rare-codon ratio, protein length, hydrophobicity, structure-derived descriptors
Evaluation
Protein-grouped holdout, RMSE / MAE / R², ablations
Explainability
Permutation importance, SHAP, residual diagnostics
Product
Multi-tab Streamlit with Plotly charts and py3Dmol structure views
Tests
26 unittest cases: training protocol safety, artifact loading, dashboard helpers, codon analysis
04

What it doesn't do (yet)

  1. The expression target is a deterministic simulated proxy built from codon, host and structure signals. It is explicitly not wet-lab measured.
  2. It ships as a CLI and Streamlit app; the REST service layer is planned, not built.

Read the code and the full results on GitHub