PhD Researcher · AI for Science · Agents

Hi, I am Encheng Su.

I am a PhD student in Artificial Intelligence at the University of Science and Technology of China, advised by Prof. Wanli Ouyang and Prof. Houqiang Li. I am currently a research intern at Shanghai AI Laboratory.

My research focuses on AI for Science and LLM agents. I build scientific reasoning benchmarks, evaluation tools, multimodal models, and long-horizon agent systems with structured memory. My broader goal is to make AI scientifically rigorous, dependable, and useful in real research workflows.

Questions that I am especially interested in:

  • How can AI reason reliably under scientific constraints, evidence, units, and processes?
  • How can agents use memory, tools, and feedback over long-horizon tasks?
  • How can multimodal systems transfer across disciplines and real-world scientific settings?

Previously, I spent one year in a research internship at UCLA (Aug 2024–Aug 2025), working as a Research Assistant with Prof. Debiao Li on longitudinal breast MRI and AI for Medicine. I also worked as a Researcher at BAAI (Mar–Sep 2025), where I focused on AI4Medical. Earlier research included 3D vision at CUHK/HKCLR, AI for Science at Microsoft Research Asia, and work with Prof. Alois Knoll at the Technical University of Munich.

Portrait of Encheng Su

News

Newest first · Scroll down for more.

S3Mem accepted to EMNLP

Structured scene–event memory for long-horizon interactive question answering.

Paper ↗

DAPD preprint released

Dual-anchored policy distillation for more stable model alignment.

Paper ↗

Structure–property reasoning preprint released

Accurate, Interdisciplinary and Transparent Structure–Property Understanding with Deep Native Structural Reasoning.

Paper ↗

CardioLens preprint released

Evaluating the clinical reality gap of MLLMs with multi-sequence cardiac MRI.

Paper ↗

Hidden in Plain Sight released

Visual-to-symbolic analytical solution inference from scientific field visualizations.

Paper ↗

SciIF preprint released

A benchmark for rigorous scientific instruction following.

Paper ↗

SciEvalKit released

An open-source evaluation toolkit for scientific general intelligence.

Paper ↗

SciReasoner preprint released

A foundation for scientific reasoning across multiple disciplines and representations.

Paper ↗

Scientific LLM survey released

A data-centric review spanning scientific data foundations, models, and agent frontiers.

Paper ↗

PhysUniBench released

Our undergraduate-level multimodal physics reasoning benchmark is now available.

Paper ↗

NIR adversarial attack preprint released

A human-imperceptible physical attack for near-infrared face recognition models.

Paper ↗

UniSTD accepted to CVPR 2025

Unified spatio-temporal learning across ten tasks and four scientific disciplines.

Paper ↗

BiSeg-SAM presented as a BIBM Oral

Weakly supervised post-processing for medical binary segmentation with SAM.

Paper ↗

Education & Experience

A concise record of my academic and research path.

Education

USTC logo
University of Science and Technology of ChinaPhD in Artificial Intelligence

Advised by Prof. Wanli Ouyang and Prof. Houqiang Li; research on LLMs and agents.

TUM logo
Technical University of MunichMaster's degree

Training in machine learning, robotics, and computer vision.

Experience

Shanghai AI Laboratory logo
Shanghai AI LaboratoryResearch Intern

Scientific intelligence, model evaluation, and agent systems.

BAAI logo
Beijing Academy of Artificial IntelligenceResearcher

AI for Medicine (AI4Medical), medical multimodal models, and evaluation.

UCLA logo
UCLAResearch Assistant

One-year research internship on longitudinal breast MRI, cancer-risk prediction, and AI for medical imaging, guided by Prof. Debiao Li.

CUHK logo
CUHK · Hong Kong Centre for Logistics RoboticsResearch Assistant

3D vision, handheld laser scanning, SDK development, and point-cloud processing.

Microsoft logo
Microsoft Research AsiaResearch Intern

AI for Science.

Midea logo
Midea GroupStudent Intern

First & Co-first Author

SciIF: Benchmarking Scientific Instruction Following Towards Rigorous Scientific Intelligence

Encheng Su, Jianyu Wu, Chen Tang, Lintao Wang, Pengze Li, Aoran Wang, Jinouwen Zhang, Yizhou Wang, Yuan Meng, Xinzhu Ma, Shixiang Tang, Houqiang Li

SciIF evaluates whether models can satisfy the explicit constraints that make scientific answers rigorous—not only produce a plausible final answer. It provides structured verification for conditions, units, assumptions, and required solution processes.

arXiv · 2026
PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level

Lintao Wang*, Encheng Su*, Jiaqi Liu, Pengze Li, Jiabei Xiao, Wenlong Zhang, Xinnan Dai, Xi Chen, Yuan Meng, Lei Bai, Wanli Ouyang, Shixiang Tang, Aoran Wang, Xinzhu Ma

* Equal contribution.

PhysUniBench evaluates conceptual, mathematical, and diagram-based reasoning across undergraduate physics. It is designed to expose failures that simpler answer-only benchmarks miss.

WAIC · 2025

Core Contributor

DAPD: Dual-Anchored Policy Distillation

Jianyu Wu, Yizhou Wang, Encheng Su, Chen Tang, Shixiang Tang

DAPD addresses the “privilege illusion” in on-policy self-distillation by aligning teacher and student behavior along two matched-information paths.

arXiv · 2026

Academic Service

Peer review and community contributions.

Best Reviewer

SIG-Cardiac Workshop on Digital Heart in the MICCAI community.

Workshop page ↗
Conference Reviewer

Reviewer for BIBM, MICCAI, AAAI, ACL, and related venues in artificial intelligence, medical imaging, and scientific machine learning.