Yi Cao News Research Experience Awards Contact Blog Talks CV GitHub Google Scholar

From Black Box to Blueprint:
Trustworthy AI for Materials Discovery

I am a PhD researcher at Johns Hopkins University developing frameworks that bridge large scale molecular dynamics with first-principles simulations by machine-learned force fields, spanning explainable AI (AAAI 2026 XAI4Science), generalization benchmarking (NeurIPS 2025 AI4Mat spotlight), and causal reasoning (KDD 2026). I have a strong foundation in both AI/ML methodology and computational materials science, with demonstrated experience translating research into industry applications.

Latest News

May 2026

Interning at Qualcomm this summer

Joined the GPU High-Level Modeling team on-site in San Diego (May 18 – Aug 23) to build an agentic, knowledge-graph-assisted causal reasoning workflow.

View LinkedIn Post →
Jul 2026

AutoMat accepted to COLM 2026

Evaluating LLM agents on scientific reproducibility in computational materials science.

Jul 2026

ARIA accepted to KDD 2026 (AI4Sciences Track)

A causal-aware framework for rescuing LLM reasoning, grown out of the 2025 LLM Hackathon Visionary Award project.

Read more See all news

Recent Work

DUAL-X

What is Your Force Field Really Learning? Gaining Scientific Intuition with a Dual-Level Explainability Framework

Yi Cao, Peter Mastracco, Jieneng Chen, Alan Yuille, Paulette Clancy*
AAAI Conference Workshop (XAI4Science), Spotlight (2026)

Dual-level explainability framework bridging model reasoning with human understanding in scientific AI.

Abstract. We present DUAL-X, a closed-loop optimization framework that integrates interpretable, human-centric rationale extraction with gradient-based attribution to give MLFF (machine-learned force field) predictions a SHAP-like audit trail — connecting atomic-position contribution scores back to chemically meaningful descriptors (SOAP, SNAP, environment descriptors).

Why it matters. The AAAI 2026 XAI4Science track spotlights methods that open the "black box" for scientific ML; DUAL-X is the first framework to ship a practical, end-to-end XAI recipe for materials MD models.

AutoMat

AutoMat: Evaluating LLM Agents on Scientific Reproducibility in Computational Materials Science

Ziyang Huang, Yi Cao, Ali K. Shargh, Jing Luo, Ruidong Mei, Mohd Zaki, Zhan Liu, William Jurayj, Somdatta Goswami, Michael Shields, Jaafar El-Awady, Paulette Clancy, William Gantt Walden, Nicholas Andrews, Benjamin Van Durme, Daniel Khashabi
Accepted to COLM Conference (2026) Accepted

Benchmark for evaluating LLM agents on scientific reproducibility, scoring both fidelity and faithfulness to the original workflow.

Abstract. AutoMat is a benchmark for evaluating LLM agents on scientific reproducibility in computational materials science. Agents receive the artifacts of a published paper and must reproduce the reported result; we score both fidelity and faithfulness to the original workflow.

Why it matters. The first benchmark to put LLM "reproducibility" under the same microscope for materials science that we already use for ML benchmarks.

ARIA

ARIA: A Causal-Aware Framework for Rescuing LLM Reasoning in Trustworthy Materials Discovery

Yi Cao, Liaoyaqi Wang, Jieneng Chen, Benjamin Van Durme, Alan Yuille, Paulette Clancy*
Accepted to KDD Conference, AI4Sciences Track (2026) Accepted

Causal-aware KG-LLM integration framework that mitigates "contextual tunneling" and improves scientific reasoning.

Abstract. ARIA (Adaptive Reasoning with Interpreted Atoms) is a causal-aware framework for rescuing LLM reasoning in trustworthy materials discovery. It combines rigid knowledge-graph retrieval, a contextual-tower contextualization module, and a tiered fallback hierarchy (Direct Match → Analogical → Fallback) to keep LLM suggestions chemically grounded.

Why it matters. Born out of the 2025 LLM Hackathon Visionary Award, ARIA is the first framework to combine KG retrieval with causal-aware hierarchical reasoning for materials science.

Migration as a Probe

Migration as a Probe: A Generalizable Benchmark Framework for Specialist vs. Generalist Machine-Learned Force Fields

Yi Cao, Paulette Clancy*
NeurIPS Conference Workshop (AI4Mat), Spotlight Talk (2025)

Benchmark that uses migration pathways as diagnostic probes to compare specialist vs. generalist MLFFs, revealing task-specific fine-tuning trade-offs.

Abstract. We use migration pathways as diagnostic probes to benchmark specialist vs. generalist machine-learned force fields (MLFFs), exposing where task-specific fine-tuning helps or hurts generalization across chemistries.

Why it matters. Spotlight talk & travel grant at NeurIPS AI4Mat 2025; a physically meaningful, generalizable probe for MLFF generalization in materials science.

CsPbBr3 self-healing defects

Low-energy Pathways Lead to Self-Healing Defects in CsPbBr3

Kumar Miskin, Yi Cao, Madaline Marland, Farhan Shaikh, David T. Moore, John Marohn, Paulette Clancy*
Phys. Chem. Chem. Phys., 27(29), 15446–15459 (2025)

Computational discovery of low-energy pathways that enable self-healing of defects in the CsPbBr3 perovskite, with implications for rational material design.

Abstract. We identify low-energy pathways by which point defects in the CsPbBr3 perovskite self-heal, pointing to design rules for more robust lead-halide perovskites.

Why it matters. Connects an atomistic self-healing mechanism to rational material design for stable perovskite devices.

Atomic Switch Control

Atomic Switch Control via Two-Mode Intercalation for Tunable 2D Materials

Yi Cao, Victor Wu, Paulette Clancy*
npj 2D Materials and Applications, under review (2025)

A two-mode intercalation mechanism that acts as an atomic switch for reversibly tuning the properties of 2D materials.

Summary. A two-mode intercalation mechanism functions as an atomic switch, reversibly modulating the interlayer structure and properties of 2D materials.

Why it matters. Offers a controllable, switchable route to tune 2D material behavior for device applications.

Experience & Skills

Technical Skills

Machine Learning & AI

  • Deep Learning (PyTorch, TensorFlow)
  • Large Language Models (fine-tuning, prompt engineering)
  • Causal Inference, Graph Neural Networks
  • Transfer Learning, Explainable AI

Scientific Computing

  • Molecular Dynamics (LAMMPS, GROMACS)
  • Density Functional Theory (Quantum ESPRESSO)
  • High-Performance Computing (MPI, CUDA)
  • Materials Informatics, Scientific AI

Programming & Tools

  • Python, MATLAB, R, Git
  • Docker, Linux/Unix
  • SQL, Database Management
  • Distributed Computing, Large-scale Data Processing

Education & Research Timeline


May 2026 - Aug 2026
Engineering Intern, GPU High-Level Modeling
Qualcomm, San Diego
Aug. 2026 - May 2028 (Expected)
MS in Computer Science
Johns Hopkins University
Aug. 2023 - May 2028 (Expected)
PhD in ChemBE
Johns Hopkins University
Nov 2023 - Present
Graduate Researcher
Clancy Lab, JHU
Jun 2024 - Jul 2024
CADD Intern
Viva Biotech
Sept 2019 - Jul 2023
B.S. Pharmaceutical Sci.
Fudan University
Dec 2022 - Feb 2023
Quality Culture Intern
Boehringer Ingelheim
Feb 2022 - Jun 2023
Undergrad Researcher
ISTBI, Fudan University
Jul 2022 - Aug 2022
Summer Research
Westlake University, Hangzhou
Aug 2021 - Dec 2021
Visiting Scholar
UC Berkeley

Work Experience

Qualcomm

Engineering Intern, GPU High-Level Modeling

May – Aug 2026  ·  Qualcomm, San Diego, CA

Developing an agentic workflow with knowledge-graph-assisted causal reasoning to automate complex debugging processes, integrating LLM capabilities with existing diagnostic tools to streamline operations.

LinkedIn Post
Viva Biotech

Computational Drug Design Intern

Jun – Jul 2024  ·  Viva Biotech, Shanghai

Conducted Computer-Aided Drug Design (CADD) research using co-solvent MD simulations, optimizing drug discovery through protein-ligand interaction analysis.

Company Website
iGEM Competition

Scientific Advisor

Dec 2022 – Nov 2023  ·  Fudan iGEM Team, Shanghai

Guided experimental design and scientific documentation. Led brainstorming sessions resulting in Gold Medal and Best Environmental Project.

View Project
Boehringer Ingelheim

Quality Culture Intern

Dec 2022 – Feb 2023  ·  Boehringer Ingelheim, Shanghai

Led a team in developing a white paper on quality culture through research and interviews, resulting in improved company-wide quality guidelines.

Company Website
Teaching

Science Communication & Teaching

Conference Talks, Posters, and More

Selected for PHM Society Doctoral Symposium 2025, presented at MRS Fall Meeting 2024, and completed JHU Teaching Institute certification.

View All

Awards & Honors

  • Best Poster Award - Women in Data Science and AI Symposium (2026)
  • Top 1 Poster Presentation Award - Women of Whiting STEM Symposium (2026)
  • 2025 Visionary Award - LLM Hackathon for Materials Science (ranked 6th of 120 teams)
  • Spotlight Talk (Top-tier recognition) - AAAI XAI4Science Workshop (2026)
  • Spotlight Talk & Travel Grant - NeurIPS AI4Mat Workshop (2025)
  • Oral Presentation, PhD Consortium - KDD 2026, in person (Aug 2026)
  • Poster Presentation Award Winner - Women in AI 2025
  • Invited Talk & Session Chair, Doctoral Symposium Selectee - PHM Society (1 of 10 PhD students globally, 2025)
  • Empower Your Pitch Finalist - JHU (1 of 12 PhD students university-wide, 2025)
  • "Graduate Star" Nomination Award - Fudan University (2023)
  • Excellent Graduates of Shanghai Colleges - (2023)
  • 1st Class Scholarship - Fudan University (2021)

Vision for AI-Accelerated Materials Discovery

Hover or click to expand; click again to collapse.

What is your long-term goal?
Build a closed-loop system linking AI + simulations + experiments.

My long-term goal is to bridge the gap between computational simulations and experimental materials science, enabling a closed-loop design process. With prior wet-lab training in biomaterials, I’ve seen how tedious trial-and-error methods are.

My vision is to build systems that minimize experiments by learning from past data and simulations—so materials discovery becomes faster, deeper, and smarter.

▲ Collapse
What is your mission as a simulation researcher?
Use ML to extract maximum insight from minimal experiments.

I aim to merge simulation data and historical experiments using advanced ML techniques like active learning and transfer learning to uncover hidden patterns. This allows us to optimize material design with fewer experiments, while gaining more knowledge—accelerating understanding of atomic-level interactions and enabling better materials in fewer cycles.

▲ Collapse
How do you understand Machine Learning?
ML is not a black box—it’s a transparent, evolving partner in science.

To me, ML is not magic—it’s a dynamic tool that gains strength when guided by domain knowledge. With strong grounding in materials science, I see ML as a transparent, explainable collaborator. CPUs and GPUs are extensions of human thought. AI and humans co-evolve, inspiring each other.

As Marie Curie once said, "Nothing in life is to be feared, it is only to be understood." Through interdisciplinary research in ML and materials, I hope to help people understand—and therefore face—the world with greater confidence and curiosity.

▲ Collapse

Get In Touch

PhD internships · research collaborations · speaking

  • Address

    3400 N. Charles Street
    Baltimore, MD 21218
    United States
  • Phone

    +1 (443) 278-3766
  • Email

    ycao73@jh.edu