Selected Projects

My research develops general machine-learning methods for scientific discovery, with molecular design and drug discovery as the primary application domain. These selected projects reflect my current focus on pretrained representations, generative models, and reward-guided optimization.

Elite-Weighted Supervised Fine-tuning (EW-SFT)

Goal-directed molecular optimization · 2026 · Biogen × UMass Amherst

EW-SFT is a reward-guided optimization method for adapting molecular generators without deriving a separate trajectory-level reinforcement-learning objective for each architecture. It maintains a rolling elite buffer of genetically evolved, high-scoring molecules and fine-tunes each generator with its native pretraining loss.

  • General method: uses reward-guided elite selection and native-loss adaptation across autoregressive, masked-diffusion, and discrete-flow molecular generators.
  • Constrained generation: evaluated on de novo design, motif extension, and linker design with 3D-shape and 2D-similarity objectives, as well as the PMO benchmark.
EW-SFT workflow: molecular generators sample candidates, genetic search evolves molecules, oracle scoring selects an elite set, and the generator is updated with an elite-weighted native loss.
EW-SFT uses a shared selection-and-native-loss update while preserving each generator's own proposal mechanism and training objective.

Pretrained Embedding Distance (PED)

Virtual screening and molecular generation · 2026 · Biogen

PED is a training-free molecular similarity measurement computed from the representations of pretrained molecular foundation models. It provides a shared representation-space measure for ligand-based virtual screening and for steering molecular generation.

  • Representation learning for screening: ranks candidate molecules using distances from pretrained molecular language, diffusion, graph-Transformer, and multimodal models.
  • Reward-guided generation: uses embedding distance as a reward signal for SMILES-based and synthesizable molecular generation, with evaluation of retrieval quality, sample efficiency, diversity, drug-likeness, and model-predicted binding affinity.
Pretrained Embedding Distance workflow connecting virtual screening and reinforcement-learning molecular generation, comparing traditional molecular similarity with distances in pretrained molecular-model embeddings.
Pretrained embedding distance connects molecular similarity, ligand-based virtual screening, and reward-guided molecular generation.

For my complete publication record, please see my Google Scholar profile.