Howdy! Iām Weijia Zhang
I am an incoming M.S. student in Computer Science at Yale University (2026 - 2028), admitted to the (Thesis Track) M.S. in Computer Science with Full Scholarship.
I graduated from UIUC in Math + Computer Science with the 2026 C.W. Gear Outstanding Undergraduate Student award, as one of two annual recipients.
Currently, I am a Research Scientist Intern at TikTok working on self-evolving agents, and a research assistant in UIUC U Lab on LLM agents, multimodal agents, and agentic RL, advised by Prof. Jiaxuan You.
Research
My research interests center on LLM agents, especially next-generation AI agents that bridge virtual and physical worlds through socially intelligent, tool-agnostic, and ethically grounded architectures.
- Multimodal agents: memory, reasoning, tool use, and multi-agent systems
- Conversational AI: anthropomorphism and social intelligence
- Post-training: agent SFT and RL
News
- 2026.06: š Joined TikTok as a Research Scientist Intern working on self-evolving agents! š¤āØ
- 2026.05: šš Graduated from UIUC in Math + Computer Science with the 2026 C.W. Gear Outstanding Undergraduate Student award! š
- 2026.01: šš Paper on bias inheritance was accepted to ACL 2026 as an Oral! šāØ
- 2025.10: š SeeingEye was released as an arXiv preprint and is under review at EMNLP 2026! šš
- 2025.07: š Joined Microsoft Research Asia as a research intern! š¬š¼
Gamedev
Beyond research, I am a passoinate indie game developer, feel free to check my game work on the game page. I am also willing to discuss the future of AI X Game.
Education
Yale University
M.S. in Computer Science
2026-09 - 2028-05
Thesis Track with Full Scholarship
GPA: 4.0/4.0
University of Illinois Urbana-Champaign
B.S. in Computer Science and Mathematics
2022-08 - 2026-05
2026 C.W. Gear Outstanding Undergraduate Student, one of two annual recipients; 2025 Dean's List
GPA: 3.7/4.0
Work Experience
TikTok
Research Scientist Intern, RSI / Self-Evolving Agents
Jun 2026 - Present
Project lead for TikTok's agent self-evolution framework, built from 0 to 1.
- Designed and shipped a closed Solver-Reflector-Evolver loop in which the Solver executes tasks, the Reflector attributes failures, and the Evolver iterates reusable skills; lifted F1 from 62.7% to 81.3% on unoriginal-video detection (9,600 real short videos) and from 76.5% to 87.3% on reposting-account detection (200 accounts).
- Designed a transferable Online Failure Discovery process: the Reflector mines failure patterns from scenarios and Solver traces, maintains a two-level scenario-to-failure-mode taxonomy, and distills high-value error clusters into reusable skills, so the same framework self-evolves across scenarios without human-predefined error types.
- Trained a video-understanding model: built an account-level SFT dataset of 74K accounts and 1.33M videos, ran distributed full-parameter SFT of Qwen3-VL 2B/4B/8B on 32 H100s with DeepSpeed, and calibrated thresholds across 53 checkpoints; the final 4B model cut false positives by 42.8% (318 to 182) against the 8B baseline at ~70% recall.
Microsoft
Research Intern, Large Language Models
Jul 2025 - Sep 2025
Built the Excel Coding Agent data engine and post-training pipeline for Excel Copilot's code agent.
- Produced execution-verified SFT/RL data by mining 20,000+ real-world workbooks, back-translating user queries from spreadsheet context, sampling multi-turn Office.js rollouts (write, execute, repair), and using SheetEngine final-state assertions as a programmatic verifier to filter incorrect and reward-hacking solutions.
- Post-trained the coding agent on this corpus with rejection-sampling SFT and RL over same-task pass/fail rollouts under execution-based reward, improving Office Scripts pass@1 by 15% against the production baseline on a held-out set.
Reborn Network
AI Agent Engineer
May 2023 - Jul 2023
Built real-time embodied role-playing agents in Unity VR.
- Built an embodied role-playing agent in a Unity VR environment, closing a real-time perception-dialogue-action loop across text, voice, and full-body VR actions at sub-second end-to-end latency.
- Introduced RAG/vector databases and dual-level (episodic + semantic) memory, improving cross-session recall accuracy from 38% to 61% on an internal multi-session dialogue eval.
my schedule
Feel free to check my availability. Times shown in Eastern Time.

