Reasoning · Learning · Multi-agent systems

Zihao Zhao 赵梓豪

Ph.D. student in Information Sciences
University of Illinois Urbana-Champaign

I’m interested in reasoning, learning, and orchestration in LLM-based multi-agent systems.

I am advised by Hao Wang at UIUC. Previously, I pursued doctoral studies at Rutgers and received my B.Eng. in Automation from Shanghai Jiao Tong University’s Zhiyuan Honors Program.

Zihao Zhao
Based at UIUCIllinois, USA
  1. NeurIPS2026
    Two-row diagram. Top: overlapping sampled agents cover the same questions, so majority vote settles on a repeated wrong answer. Bottom: diverse agents cover complementary questions, and after rounds of Latent Verification Debate the readout recovers the correct minority answer.

    When Debate Helps: Proposal Supply and Verification-Aware Readout in Multi-Agent Reasoning

    Zihao Zhao, Tunyu Zhang, Haizhou Shi, Yusong Zhao, Xinxi Zhang, and Hao Wang

    NeurIPS 2026 · Accepted

    Figure 1 The supply–readout view: diverse agents supply complementary proposals, and verification-aware readout (LVD) recovers a correct minority answer that majority voting misses.

  2. NeurIPS2026
    Two plots of accuracy against total tokens used on AIME 24&25. Left: for multi-agent test-time scaling, Orch-RM reaches the highest accuracy with far fewer tokens than LLM-judge, log-probability, and majority-vote baselines. Right: for continued orchestrator training, Orch-RM reaches 68.33% accuracy, about 3 points above GRPO, DPO, and RFT, while using roughly ten times fewer tokens than GRPO.

    Reward Modeling for Multi-Agent Orchestration

    King Yeung Tsang†, Zihao Zhao†, Vishal Venkataramani, Haizhou Shi, Zixuan Ke, Semih Yavuz, Shafiq Joty, and Hao Wang

    NeurIPS 2026 · Accepted

    Figure 1 OrchRM improves the accuracy–token trade-off for multi-agent test-time scaling (left) and continued orchestrator training (right).

  3. NeurIPS2026
    Two-panel diagram. Panel A: a stable and an unstable three-round agent interaction reach the same final answer but deserve low and high uncertainty, which a static view of the final answers cannot distinguish. Panel B: SAUCE turns each round's outputs into observations and sequentially updates reliability-weighted beliefs to produce an uncertainty score.

    Sequential Probabilistic Uncertainty Estimation for Parallel Multi-Agent Reasoning Systems

    Tunyu Zhang, Zihao Zhao, Yusong Zhao, Haizhou Shi, Zhuohang Li, Haoxian Chen, Hao Wang, and Dimitris N. Metaxas

    NeurIPS 2026 · Accepted

    Figure 1 Uncertainty depends on how agents interact across rounds, not just on their final answers; SAUCE updates uncertainty round by round.

† Equal contribution. My name is shown in bold.

Along the way

Background

Full curriculum vitae

Education

Sept. 2026–Present

University of Illinois Urbana-Champaign

Ph.D. student in Information Sciences
Advisor: Hao Wang

Aug. 2025–Aug. 2026

Rutgers University–New Brunswick

Doctoral studies in Computer Science
Advisor: Hao Wang

Sept. 2021–June 2025

Shanghai Jiao Tong University

B.Eng. in Automation
Zhiyuan Honors Program

Academic service

Reviewer for ACL Rolling Review (2025–2026) and NeurIPS (2026).

Teaching

Teaching assistant for Database at UIUC, and Computer Architecture and Introduction to Deep Learning at Rutgers.

Beyond research

Previously a senior student network administrator on the SJTU-NIMO core team, supporting campus connectivity and mentoring junior administrators.

Selected recognition

National First Prize, TI Cup China Undergraduate Electronics Design Contest, 2023.

Let’s connect

Interested in multi-agent systems?

I welcome conversations about research and opportunities in academia and industry.

zihaoz@illinois.edu