Protein Plant Synthesis
If you had to produce a therapeutic protein in a plant, which host and which DNA design should you try first?
A leakage-controlled surrogate ML pipeline that designs DNA for therapeutic proteins across four plant hosts, and explains its own predictions.
From UniProt to a ranked decision
Proteins are fetched from UniProt and quality-checked. Each one gets a naive back-translation and host-aware optimised DNA for four plant hosts, using host-biased codon tables. Structure descriptors come from AlphaFold, with a deterministic fallback so the app still runs offline.
Codon adaptation (CAI), GC compatibility, rare-codon ratios and protein features feed a set of regressors. The final screen turns predictions into an operations view: which host to try, at what estimated cost and time.

Making the evaluation credible
Grouped splits
The same protein appears once per host. Random splits would leak an accession into both train and test, so splits are grouped by protein (GroupShuffleSplit, GroupKFold).
Beat the trivial model first
Every model is compared against a dummy baseline that predicts the mean. The dashboard asks the question explicitly: is this better than a trivial prediction?
Explain, then trust
Permutation importance, local SHAP explanations and residual analysis are part of the product, not an afterthought.




One more section for engineers: architecture, hyper-parameters and design decisions.
Under the hood
Under the hood- Models
- Dummy baseline, Linear, Ridge, Random Forest (350 trees, depth 14), XGBoost (350, depth 6, lr 0.05), optional LightGBM
- Features
- CAI, GC content, codon diversity, rare-codon ratio, protein length, hydrophobicity, structure-derived descriptors
- Evaluation
- Protein-grouped holdout, RMSE / MAE / R², ablations
- Explainability
- Permutation importance, SHAP, residual diagnostics
- Product
- Multi-tab Streamlit with Plotly charts and py3Dmol structure views
- Tests
- 26 unittest cases: training protocol safety, artifact loading, dashboard helpers, codon analysis
What it doesn't do (yet)
- The expression target is a deterministic simulated proxy built from codon, host and structure signals. It is explicitly not wet-lab measured.
- It ships as a CLI and Streamlit app; the REST service layer is planned, not built.