Research Portfolio

Publications

Research on reliable AI agents, multimodal reasoning, and the social and ethical behavior of large language models.

Google Scholar
Abstract visualization of research papers connected through a knowledge graph

Latest Research

2026

Agent reliability, multimodal reasoning, and responsible AI.
First Author · Under Review

How Much Vision Does Multimodal Reasoning Need? Vision-Stripping for Multimodal Benchmarks

Under review
Vision-stripping levels for multimodal benchmark reasoning

Weijia Zhang, Zijia Liu, Tianyi Zhang, Ruiqi Chen, Lian Zhang, Haoru Li, Haoqi Chen, and Jiaxuan You

Vision-stripping for multimodal benchmarks and multimodal reasoning.

  • Multimodal Reasoning
  • Vision Stripping
  • Benchmarks
  • VLM Evaluation
First Author

Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs

EMNLP 2026
Single-scene moral calibration and composite judgment framework

Weijia Zhang*, Ruiqi Chen*, Yunze Xiao*, and Weihao Xuan

A study of compressed moral composition in frontier LLMs.

  • LLM Ethics
  • Moral Reasoning
  • Frontier Models
  • Evaluation
First Author

CUADebug: Diagnosing and Repairing Computer-Use Agent Failures

EMNLP 2026
CUADebug pipeline from failed trajectory diagnosis to debugger-guided rerollout

Weijia Zhang, Kunlun Zhu, Zeyi Liu, Yinting Chen, Tianyi Ma, Jiateng Liu, Jiaxun Zhang, Bingxuan Li, Pan Lu, Xiangru Tang, Heng Ji, and Jiaxuan You

A framework for diagnosing and repairing computer-use agent failures with a CUA-specific error taxonomy, benchmark, and tool-augmented debugger.

  • Computer-Use Agents
  • Failure Diagnosis
  • Self-Evolution
  • Agent Memory
ACL Oral

Understanding and Mitigating the Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks

ACL 2026, Oral
Bias inheritance research pipeline from synthetic data to downstream mitigation

Miaomiao Li, Hao Chen, Yang Wang, Tingyuan Zhu, Weijia Zhang, Kaijie Zhu, Kam-Fai Wong, and Jindong Wang

A systematic study of how bias can be inherited through LLM-generated synthetic data and how mitigation strategies behave across tasks.

  • Bias Inheritance
  • Synthetic Data
  • Data Augmentation
  • Mitigation

Research in 2025

2025

Agentic multimodal reasoning and learning from agent failures.
Co-first Author · Under Review

SeeingEye: Agentic Information Flow Unlocks Multimodal Reasoning in Text-only LLMs

arXiv:2510.25092 · AAAI 2026 under review
SeeingEye translator and text reasoning agent information flow

Weijia Zhang*, Zijia Liu*, Haoru Li*, Haoqi Chen*, and Jiaxuan You

Agentic information flow for unlocking multimodal reasoning in text-only LLMs.

  • Agentic Information Flow
  • Text-only LLMs
  • Multimodal Reasoning
  • Tool Use
ICML Workshop

Where LLM Agents Fail and How They Can Learn From Failures

ICML 2026 Workshop
AgentDebug robot representing memory, planning, reflection, and action

Kunlun Zhu, Zijia Liu, Bingxuan Li, Muxin Tian, Yingxuan Yang, Jiaxun Zhang, Pengrui Han, Qipeng Xie, Fuyang Cui, Weijia Zhang, Xiaoteng Ma, Xiaodong Yu, Gowtham Ramesh, Jialian Wu, Zicheng Liu, Pan Lu, James Zou, and Jiaxuan You

A study of LLM agent failures and how agents can learn from failed trajectories.

  • LLM Agents
  • Failure Modes
  • Learning from Failure
  • Agent Evaluation

Research in 2024

2024

Emergent social behavior in multi-agent systems.
Co-first Author

Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory

The First Workshop on AI Behavioral Science, ACM SIGKDD 2024
Artificial Leviathan sandbox society simulation environment

Gordon Dai*, Weijia Zhang*, Jinhan Li, Siqi Yang, Srihas Rao, Arthur Caetano, and Misha Sra

A simulated LLM agent society for studying emergent social contracts through Hobbesian social contract theory.

  • Multi-Agent Systems
  • Social Simulation
  • LLM Agents
  • AI Behavior