Dry Lab · Software
Software
Our software is meant to be accessible on phones, tablets, and other devices
How Our Software Works
Our software is designed to help us find and evaluate potential improvements to the enzymes used in our PETosome system. Instead of testing thousands of possible mutations experimentally, we use computational tools to narrow down the possibilities and identify candidates that are worth testing in the laboratory.
The process starts with an existing protein sequence. From there, our software uses a combination of protein language modeling, structure prediction, and structural analysis to evaluate how different changes could affect the protein.
1. Starting With a Protein
The process begins with the amino acid sequence of one of our enzymes, such as PETase or MHETase. This sequence is used as the starting point for generating possible protein variants.
We can either analyze the original protein or introduce changes at specific positions that we think could affect its function.
2. Generating Possible Mutations
The next step is to explore possible changes to the protein sequence.
Using ESM2, a protein language model, we can evaluate different amino acid substitutions and determine how reasonable each sequence is based on patterns learned from millions of protein sequences.
This allows us to explore mutations more systematically instead of choosing them entirely by hand.
For example, if we are interested in a specific region around an enzyme's active site, we can generate different mutations at the surrounding positions and compare them.
3. Predicting Protein Structure
After generating potential sequences, we need to understand what those mutations could do to the protein's structure.
Our software uses ESMFold to predict the three-dimensional structure of each candidate.
This gives us a way to compare the predicted structures of the original protein and the mutated versions.
4. Analyzing the Structure
The predicted structures are then analyzed using our own computational methods.
We focus on structural features that are relevant to enzyme function. For example, when working with MHETase, we can examine the region around the substrate-binding pocket and compare how different mutations change this area.
For PETase, we can also investigate how changing the linker connecting the enzyme to the dockerin could affect the position of the dockerin relative to the active site.
This allows us to connect the predicted structure back to the actual problems we observed in our experiments.
5. Scoring and Comparing Candidates
Once the candidates have been analyzed, the software combines the information into scores that allow us to compare different designs.
Instead of looking at hundreds of structures individually, we can use these scores to identify the candidates that appear most promising.
The goal is not for the software to tell us that one mutation is definitely better than another. Instead, it helps us prioritize which candidates should be investigated further.
6. Selecting Candidates for the Wet Lab
The most promising candidates can then be selected for experimental testing.
This is where our computational and lab work connect.
The software helps reduce the number of possible designs we need to consider, while experiments allow us to determine whether the computational predictions actually translate into improved enzyme performance.
The results from the experiments can then be used to evaluate and improve our computational approach.
The Complete Process
Overall, our software follows a cycle:
- Protein Sequence
- Generate Possible Mutations
- Evaluate Sequences With ESM2
- Predict Structures With ESMFold
- Analyze Structural Features
- Score & Compare Candidates
- Select Promising Designs
- Experimental Testing (All our data is theoretical and needs real world testing to prove reliability)
Rather than replacing experimental testing, our software helps us decide what is worth testing in the first place. By reducing the number of possibilities and giving us a way to compare different designs, we hope to make the process of engineering our PETosome enzymes more efficient and more informed.
Software Engineering Efforts
This page documents the efforts expended on software engineering, some of which resulted in running software.
- Use of OpenFold was our initial emphasis, but we failed to recruit a gene folding expert to interpret output. We did not add code to OpenFold.
- Between January and June, most of our software attempted to use game-like interfaces to teach students about the mathematical models that we believed should apply to our experimental results. The game arcade helped acclimate the students at the start of the project.
- Animations for the website were made that falls under the Creative Commons license.
- After June, we developed further Python/Jupyter software to predict useful directions for future research.
The PETosome AI Design Toolkit
Alongside the mathematical models above, we built a small AI pipeline to help diagnose and fix two problems that showed up once our fusion enzymes reached the wet lab. It combines ESM2 (a protein language model, used to score how biologically plausible a sequence is) with ESMFold (structure prediction, used to check whether a designed sequence still folds sensibly and to measure distances/pocket volumes on the predicted structure).
The two problems it addresses
1. The dockerin blocks the active site. Fusing a dockerin “hook” onto our PETase (ICCG–DoT) to attach it to the PETosome scaffold cut its activity roughly 6× on a small-molecule assay, and TPA production on real PET film dropped progressively with reaction depth (7% for BHET, 29% for MHET, 44% for TPA), a signature of the dockerin physically crowding the enzyme's “doorway.” We used ESM2 + ESMFold to design and score alternative linkers that push the dockerin further from the active site.
2. MHETase is the bottleneck. Even without any fusion, our PETase produces MHET faster than our MHETase (TfCa–DoG) can consume it: after 96 hours, 67% of product was still stuck as MHET and only 32% had reached the final product, TPA. We had already rationally designed one improved mutant (TfCaWA, carrying I69W/V376A) to widen the substrate pocket, but hadn't tested it yet or asked whether a better substitution exists. We used ESM2 to scan every pocket-lining residue for plausible substitutions, then predicted structures for the resulting candidate mutants.
Pipeline & environment
Both notebooks run in a reproducibility-pinned Conda environment (Python 3.10, PyTorch 2.12.1 + CUDA 12.1, Transformers 5.12.1, BioPython 1.87, py3Dmol for visualization), with random seeds fixed to 42 across NumPy, PyTorch, and Python's random so that results are identical across machines running the same model weights. Each notebook ranks its candidates with a composite score blending three signals: ESM2 sequence plausibility, ESMFold pLDDT (structure-confidence, >70 considered reliable), and a geometry term (active-site distance for linkers; pocket volume for mutants). The run below is the real model (ESM2–650M / ESMFold v1, on an RTX 3060 12GB), independently verified: every structure file was checked for genuine multi-atom-per-residue geometry and a fresh generation timestamp, ruling out leftover placeholder data.
Linker optimization: predicted structures
Nine candidate linkers (five rational glycine–serine repeats, four AI-guided sequences) were generated and scored; the full ranking table and comparison chart are archived on the Raw Data page. GS_len25 tops the composite score, but as the caveat above explains, its real active-site distance is actually shorter than the native linker's, so this is not yet a confident pick:


We're treating the composite ranking as provisional until the distance discrepancy is resolved; see Raw Data for the full scoring table across all four retained candidates.
MHETase pocket engineering: predicted structures
ESM2 scanned 11 residues lining the MHETase substrate-binding pocket; suggestions were combined into 8 candidate mutants (full ranking on Raw Data). Position 69 comes up as the most-suggested site to mutate, consistent with our original rational design (TfCaWA, I69W/V376A), and ESM2 favors phenylalanine (I69F) there over the tryptophan we already picked, but as the caveat above explains, the pocket-volume metric can't currently distinguish any of these constructs, so this shouldn't be read as a ranked recommendation yet:





The I69F preference is still a testable hypothesis worth pursuing (it's independently consistent with our own rational design), but we're not treating the composite ranking as evidence either way until pocket_free_volume() is fixed to account for side-chain identity and the two hardcoded control scores are replaced with real ESM2 output. We plan to synthesize and test I69F and I69F/V376A side-by-side with TfCaWA regardless, in the BHET kinetics assay. Full rankings on Raw Data.
Full executed notebooks & reports
The complete, executed analysis, every cell, figure, and intermediate table, is embedded below, along with the plain-language explainer and the verification report that surfaced the ranking issues described above.
Verification report · confirms the run is genuine model output and diagnoses the two ranking issues above
Open: PETosome verification report (static/reports/petosome_verification_report.html)
Full explainer · plain-language walkthrough of the biology, the tools, and what to do next
Open: PETosome full explainer (static/reports/petosome_full_explainer.html)
01_linker_optimization.ipynb · Linker design report
Open: Linker optimization notebook (static/reports/01_linker_optimization.html)
02_mhetase_optimization.ipynb · MHETase engineering report
Open: MHETase optimization notebook (static/reports/02_mhetase_optimization.html)
Open Source Licenses
Our scientific software is made available under an OSI-approved open-source Apache license.
Generative AI and Large Language Models
ESM2 and ESMFold to speculate about future experiments.
GPT-5-Codex for animation, Gemini Flash 5 for image generation, and Claude Sonnet for project planning.
GPT 5.6 Sol to generate images.
Google Gemini 3.5 was used to search for website links, but the websites were examined by humans and summarized by humans. Early brainstorming ran human-generated ideas through Google Gemini 3.5 to start discussions that were later debated and hand-written by humans, so no wiki documents are the outputs resulting from those generations. The lecture content was later completely revised by teams of humans so that no Gemini content can be identified.
JavaScript, CSS, and HTML were initially copied from iGEM templates and heavily modified by humans. Claude Sonnet 5 was used to fix extensive bugs introduced by human changes. Claude Sonnet 5 incidentally re-wrote some of the prose, but as much prose as possible was identified and manually rewritten by humans.