Shuhao Zhang

About

Building agents that understand the world they act in, decide with that understanding, and improve themselves with it.

I'm Shuhao Zhang, an M.S. student in Computer Science and Engineering at UC San Diego, advised by Prof. Pengtao Xie, Associate Professor of ECE and co-founder of AIBuildAI. I work in his lab as a Research Assistant (UC GSR) on AI agents and LLM reasoning. I also collaborate with Prof. Yujia Zheng, Assistant Professor of Statistics and CS (affiliate) at UIUC, on causality and representation learning.

I work along three directions:

  • Turn data into readable knowledge. Find what each unnamed variable means from the causal role it plays, so that measurements become knowledge a person can read and check.Project: CausalBridge (ICLR 2027, under review)
  • Predict before you run. Estimate how accurate and how costly each model or tool will be on an input before running it, and use the estimate to decide what to run.Projects: SCOPE (LLM routing, ICML 2026), SCOUT (prompt-injection defense, EMNLP 2026)
  • Act on what a model knows about itself. Read a model's confidence before and after it solves a problem, and use it to accept, retry, or aggregate answers at test time.Project: Metacognitive Harness
Interests
  • Agent Systems
  • LLM Reasoning
  • Causal Representation Learning
  • World Models
  • Self-evolving Systems
  • Reinforcement Learning

I'm currently looking for PhD positions starting Fall 2027.

News

  • 2026-10Serving as a reviewer for ICLR 2027.
  • 2026-09Started as a Research Assistant (UC GSR) in Prof. Pengtao Xie's lab at UC San Diego.
  • 2026-09Submitted How Causality Bridges the Semantic Gap (sole first author) to ICLR 2027.
  • 2026-08Send a SCOUT First is accepted to the EMNLP 2026 Main Conference.
  • 2026-07On-site moderator at the 3rd FM4LS workshop, ICML 2026, in Seoul. Grateful to Prof. Pengtao Xie for the opportunity.
  • 2026-05Joined Prof. Yujia Zheng's group at UIUC as a Visiting Research Intern.
  • 2026-04Models Under SCOPE is accepted to ICML 2026.

Publications

† equal contribution

Figure 2: real systems with hidden variable names, the three steps of structure learning, embedding optimization, and translation, and the recovered names compared with the strongest baseline

ICLR 2027 · Under Review

How Causality Bridges the Semantic Gap

Shuhao Zhang, Xuran Zhou, Han Guo, Pengtao Xie, Yujia Zheng

Measurements record how a system behaves, yet they seldom say what its variables mean. CausalBridge reads the meaning of each unnamed variable from the causal role that it plays, and names observed and latent variables in questionnaires and robot systems more accurately than methods that rely on association.

Figure 1: detectors behave differently, SCOUT predicts and triages per request, and one threshold sets the safety-utility trade-off

EMNLP 2026 · Main

Send a SCOUT First: Pre-hoc Reasoning for Adaptive Detector Allocation in Prompt-Injection Defense

Shuhao Zhang†, Jiarui Li†, Qi Cao, Ruiyi Zhang, Pengtao Xie

Every prompt-injection detector stops some attacks and misses others. SCOUT predicts each detector's reliability and latency from how it behaved on similar inputs, then decides per request which detectors to run and whether to call an LLM judge.

  • −46% attack success
  • −40% wall-clock

Both against an always-on GPT-4o judge, at a 5.1-point benign-utility cost.

Figure 2: fingerprint construction, performance prediction by SCOPE, and the budget-aware decision

ICML 2026

Models Under SCOPE: Scalable and Controllable Routing via Pre-hoc Reasoning

Qi Cao†, Shuhao Zhang†, Ruizhe Zhou, Ruiyi Zhang, Peijia Qin, Pengtao Xie

A service with many LLMs must choose one model for each query. SCOPE predicts the accuracy and the cost of every candidate model before any of them runs, so the router can trade cost for accuracy and handle models it has never seen.

  • up to 95.1% lower cost
  • up to 25.7% higher accuracy
Figure 1: feeling of knowing and judgment of learning track correctness but do not guide reasoning effort

Preprint 2026

LLMs Know When They Know, but Do Not Act on It: A Metacognitive Harness for Test-time Scaling

Qi Cao, Yufan Wang, Peijia Qin, Shuhao Zhang, Pengtao Xie

Strong LLMs can tell whether they know an answer, yet they do not act on it. The harness asks a frozen model for a feeling of knowing before it solves a problem and a judgment of learning after, and uses the two signals to accept, retry, or aggregate.

  • pooled accuracy 48.3 → 56.9
  • no parameter updates

Earlier work

  • IDC-CDR: Cross-domain Recommendation based on Intent Disentanglement and Contrast Learning. Jing Xu, Mingxin Gan, Hang Zhang, Shuhao Zhang. Information Processing and Management, 2024.Advisor: Prof. Mingxin Gan
  • Using Machine Learning Models for Short-Term Prediction of Dissolved Oxygen in a Microtidal Estuary. Mina Gachloo, Qianqian Liu, Yang Song, Guozhi Wang, Shuhao Zhang, Nathan Hall. Water 16(14):1998, 2024.Advisor: Prof. Yang Song

Research Experience

  • Fall 2026 – now

    Research Assistant (UC GSR), PXie Lab, UC San Diego

    PI Prof. Pengtao Xie, Associate Professor of ECE. Research Intern in the same lab from Oct 2025. Lead of SCOUT, co-lead of SCOPE.

  • May 2026 – now

    Visiting Research Intern, Yujia Zheng’s Group, UIUC

    Advisor Prof. Yujia Zheng, Assistant Professor of Statistics and CS (affiliate). Project leader of CausalBridge.

Industry Experience

  • Jun – Sep 2026

    Research Intern, AI Agents, AIBuildAI Inc.

    Robot world models for an autonomous ML-engineering agent. In the public RT-1 study the agent post-trained a 2B NVIDIA Cosmos world model in 29 hours with no human intervention (PSNR 17.75 → 25.56). Blog

  • Jul – Aug 2024

    Algorithm Intern, Shanghai Huapu Zhihui Automation System Engineering Co., Ltd.

    Simulation and AGV scheduling for an airport warehouse system.

Education

  • 2025 – 2027

    M.S. in Computer Science and Engineering, UC San Diego

  • 2021 – 2025

    Bachelor's degree, University of Science and Technology Beijing

  • Fall 2023

    Visiting Student, UC San Diego

Academic Service

  • Fall 2026

    Reviewer, ICLR 2027

  • Jul 2026

    On-site Session Moderator, 3rd Workshop on Multi-modal Foundation Models and Large Language Models for Life Sciences, ICML 2026, Seoul

    Ran the full-day program: seven invited talks, a panel, and two poster sessions. ICML page · Schedule