Research

Research

Making the invisible steps of natural product biosynthesis visible

Overview

Overview

Much of what we eat, wear and build with ultimately traces back to natural products made by plants and microbes. Terpenoids in particular underpin pharmaceuticals, agrochemicals, fragrances, adhesives, polymer additives and electronic materials.

The remarkable skeletons of these molecules are assembled inside enzymes by complex, multi-step cascade reactions that proceed quickly and with exquisite selectivity — transformations that remain difficult for synthetic organic chemistry.

That very efficiency means the pathways are crowded with intermediates and transition states that are highly reactive, extremely short-lived and impossible to isolate. Experiment alone therefore cannot fully resolve how these reactions work.

This is where computational chemistry takes the lead. Quantum chemical calculations give direct access to the structures and energies of species that experiment cannot capture, while molecular dynamics and machine learning reveal how the enzyme active site — the reaction field itself — selects one product out of many possibilities.

Understanding, and ultimately redesigning, nature's synthetic logic is valuable not only as fundamental chemistry but also as a practical toolbox for drug discovery and materials science.

Overview of natural product biosynthesis research
01

Accelerating calculations with machine-learning potentials

The single biggest obstacle to studying reaction mechanisms computationally has been cost. Locating a transition state with quantum chemistry can take days for one reaction — far too slow to compare every pathway worth considering.

We therefore combined UMA (Universal Model for Atoms), a general-purpose machine-learning potential trained on the 120-million-molecule OMol25 dataset, with reaction-path searching by the Direct MaxFlux (DMF) method. With this framework, a mechanistic analysis that used to take days now finishes in about four minutes per reaction on average.

Speed alone would not be enough, but the accuracy holds up: across a benchmark of organic reactions, the transition-state structures had a mean RMSD of 0.24 Å and the activation energies a mean absolute error of 1.65 kcal/mol.

What this buys is not merely less waiting. An analysis that could only ever confirm one favoured pathway becomes one that can enumerate and compare every plausible pathway. The quantum chemistry and enzyme simulations described below rest on this foundation.

Reaction path search accelerated by a machine-learning potential

Chem-Station article (in Japanese)

02

Reading mechanisms with quantum chemistry

Terpene biosynthesis is a cascade in which a carbocation, generated from a linear precursor, undergoes successive cyclizations, hydride shifts and alkyl migrations. The intermediates are so short-lived that direct experimental observation is essentially impossible.

Using density functional theory (DFT), we compute each cationic intermediate and transition state and map the full energy landscape. Comparing these landscapes with experimental and isotope-labelling data lets us test proposed mechanisms — and sometimes propose entirely new pathways.

Calculations are performed with quantum chemistry programs such as Gaussian, with conformational searches carried out using tools such as Crest. In-house workflows automate and parallelize the exploration of large numbers of candidate pathways.

Reaction pathway analysis by quantum chemistry
03

Seeing enzymes through simulation and AI

Starting from the same substrate, different enzymes give different products. What makes the difference is the shape and dynamics of the active site that wraps around the substrate. We follow the motion of enzyme–substrate complexes with molecular dynamics (MD), and treat the reactive region quantum mechanically with QM/MM, in order to understand how the enzyme — as a reaction field — steers the cascade.

More recently we have been incorporating AI techniques such as machine-learning potentials and structure-prediction models, opening up large-scale pathway searches and reaction dynamics that were previously out of reach. With Rosetta we also pursue the computational design of artificial enzymes that deliver a desired skeleton.

Software we mainly use: Gaussian / Amber / GROMACS / Rosetta / Conflex

04

In concert with experiment

A computational hypothesis only becomes meaningful once it has been tested. We collaborate closely with experimental groups in Japan and abroad — synthesising authentic standards, probing reactions, and preparing and assaying mutant enzymes by genetic engineering — to put our predictions to the test.

Conversely, an experimental result that resists explanation is the best possible starting point for calculation. Moving back and forth between theory and experiment is how we work.

Collaboration with experimental science

Related papers

See all publications

Join us

For prospective students and collaborators

For details about the group, see the Computational Biology Laboratory website. For collaboration enquiries please write to hsato[at]bi.a.u-tokyo.ac.jp (replace [at] with @).