Publications
2026
arXiv
A benchmark of 169 expert-curated molecular dynamics simulation tasks that exposes fundamental
limitations of current AI agents in scientific computing — the best agent solves only 21% of easy-level
tasks.
2026
arXiv
A position paper identifying four fundamental limitations of current agentic AI scientists — problem
selection bias, LLM training gaps in laboratory practice, loss of output diversity, and benchmarks that
lack physical feedback loops — that prevent true autonomous scientific discovery.
2024
ECAI 2024
A framework that uses calibrated AI model probabilities to identify cost-effective human-AI
collaboration strategies for classification tasks, evaluated on CIFAR-10H and ImageNet-16H.
Talks
July, 2026
Building AI Agents from Scratch
July, 2023
Database Selection Tool: Empowering Decision-Making for Optimal Performance