Xuan JiangHigh performance computing and agents

New: a weekly AI newsletter purely for fun and making friends Please subscribe, no pledge

Make compute fast.Let agents learn.

I'm Xuan Jiang, a UC Berkeley Ph.D. and Research Affiliate at MIT. I've been in AI since 2019, working on mixture-of-experts models end to end, from fine-tuning and compression to the GPU communication they need at scale. I also built self-evolving agents with the reef team and co-authored CORAL.

Photo of Xuan Jiang
TimelineFrom GPU traffic simulation to AI infrastructure and agents
  1. Tongji UniversityB.S., 2016 to 2020
  2. Shanghai HuimiceAI Scientist, 2019 to 2020
  3. UC BerkeleyM.S. and Ph.D., 2020 to 2024
  4. Berkeley LabGraduate Student Researcher, 2021 to 2023
  5. MITResearch Affiliate, 2023 to present
  6. GoogleSoftware Engineer, 2024 to 2026
  7. AWS Annapurna LabsSr. HPC/AI Engineer, 2026 to present
7
Years in AI
394
Google Scholar citations
11
h-index
26
Papers and preprints
30
Merged reef/CORAL PRs

Google Scholar and GitHub, October 2026. AI experience counted from 2019.

Publication venuesICLRICMLCOLMEMNLPTransportation Research Part CIEEE T-ITSACM SIGSIM PADSTransportation Research RecordInformation FusionJournal of Air TransportationICRAT

Figures from my papers

Research

From models to systems.

  1. MoE and efficient fine-tuning

    Capabilities in MoE models concentrate in a few FFN intermediate dimensions, and rarely activated, long-tailed experts retain useful knowledge. I work on preserving both during compression or fine-tuning, and on fine-tuning that uses less GPU memory.

  2. GPU communication

    Each MoE layer dispatches tokens to experts on other GPUs and combines their results. I write kernels and transport code for this path, including kernels for transports without ordered delivery.

  3. Self-evolving agents

    I build agents that use their own records and feedback to keep improving. Updates can change model weights or revise the prompts and rules they follow.

  4. HPC and simulation

    My systems work began with simulation. I built a simulator that runs city-scale traffic across several GPUs, with applications in urban and air mobility.

Open source

Selected researchand open source.

Multi-agent research, COLM 2026

CORAL

CORAL lets long-running agents explore, reflect, and collaborate through shared persistent memory and asynchronous execution. The paper reports new state-of-the-art results on 10 tasks.

Read the paper
13631103

Four co-evolving CORAL agents set a new record on Anthropic's kernel engineering task, measured in cycles where lower is better.

  1. Simulation system, Transportation Research Part C 2024 LPSim I built LPSim for traffic simulation across multiple GPUs, using graph partitioning and ghost-zone vehicle migration. It handles networks with over 223,000 nodes in seconds on one GPU. I also worked with James Demmel and Roofline model creator Samuel Williams on in-depth performance optimization. View the code
  2. Model fine-tuning, ICLR 2025 Sparse Matrix Tuning SMT updates only the most significant sub-matrices selected from gradient updates. It outperforms LoRA and DoRA across tasks on LLaMA-family models, using 67% less GPU memory than full fine-tuning. Read the paper
  3. Reinforcement learning, ICML 2026 RAST-MoE-RL RAST-MoE-RL, a 12M-parameter model, uses regime-aware modeling and a self-attention mixture-of-experts encoder to set pre-match waits for batched ride-hailing requests and vehicles, cutting average matching delay by 10% and pickup delay by 15% on real San Francisco trajectories. Read the paper
  4. Air mobility, ICRAT 2024Best Paper Award Wind-Resilient AAM Networks Our simulation-optimization framework models wind variability and nonlinear charging across multiple vertiports to optimize advanced air mobility (AAM) fleet size and scheduling. Wind affects fleet size even on short flights, with larger effects on longer ones. Read the paper
Open-source contributions

GPU communicationfor MoE.

These are my open-source contributions to communication for mixture-of-experts (MoE) models, submitted under my own name.

  1. DeepEPdeepseek-ai/DeepEP My pull request adds opt-in dispatch and combine kernels that need only weak signals, enabling hybrid mode on transports without ordered delivery. The EFA version is open source in amazon-contributing/DeepEP, and my awsome-distributed-ai recipe builds it on EFA clusters. Open pull requestPull request #732EFA repositoryEFA build recipe
  2. NVSHMEMNVIDIA/nvshmem Worked with the NVSHMEM team on libfabric transport: RMA batching, signal-path deadlock and hang fixes, completion-queue draining, and pre-release testing for 3.7.0. Released in NVSHMEM 3.7.0My commits
  3. vLLMvllm-project/vllm Fixes host heap corruption with NCCL 2.31 and later. MergedPull request #53008
  4. rdmatopuccl-project/rdmatop Adds headless CSV recording of RDMA counters at sub-millisecond sampling intervals. MergedPull request #48
Self-evolving agents

Agents learnfrom what they do.

I contribute to reef and CORAL, two open-source systems for self-evolving agents. reef uses interaction records and feedback to update model weights or the agent's harness: prompts, rules, and skills. CORAL runs long-lived agents that explore and collaborate through shared persistent memory.

The reef learning loopBased on the learning loop documented by reef.
  1. ServeServe requests and record interactions.
  2. ObserveMatch feedback to recorded interactions.
  3. GrowProduce updates from eligible records.
  4. CommitEvaluate candidates and publish accepted updates to version history.
Open source, team member

reef

reef connects agent inference, feedback, learning, and versioned delivery so agents can continually improve.

I'm on the team, with 25 PRs authored and 18 merged.

I added reef's Claude Code harness adapter and CORAL test-time-training example, and fixed training-stack bugs. For SAO, an asynchronous RL method from another group's paper, I added the paper's value-model warmup and learning rate to reef's example, plus batch-128 runs configured like the paper's, documenting results and limitations.

Open source, COLM 2026

CORAL

CORAL is an open-source research system where autonomous coding agents explore and collaborate through shared persistent memory.

I co-authored the COLM 2026 paper and authored 14 PRs, with 12 merged.

I updated the LLM gateway to record streamed responses in request logs and added a hook for supplying headers. I also fixed CLI and agent runtime issues and updated the docs.

Weekly

I write for funand to make friends.

I'll write one post a week on Substack about interesting AI ideas and applications, plus research. Subscribers will get each post by email.

Subscribe on Substack Visit the newsletter

It's free, with no pledge. Unsubscribe any time.

Publications

Published papersand preprints.

Papers are listed newest first. Filter by topic or show only selected work, and visit Google Scholar for the complete record of papers and preprints.

  1. CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery

    Ao Qu*, Han Zheng*, Zijian Zhou*, Yihao Yan, Yihong Tang, Shao Yong Ong, Fenglu Hong, Kaichen Zhou, Chonghe Jiang, Minwei Kong, Jiacheng Zhu, Xuan Jiang, Sirui Li, Cathy Wu, Bryan Kian Hsiang Low, Jinhua Zhao, Paul Pu Liang

    Conference on Language Modeling (COLM), 2026

    51 citations
  2. Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning

    Haoze He, Xingyuan Ding, Xuan Jiang, Xinkai Zou, Alex Cheng, Yibo Zhao, Juncheng Billy Li, Heather Miller

    Conference on Language Modeling (COLM), 2026

    1 citation
  3. Less is MoE: Trimming Experts in Domain-Specialist Language Models

    Haoze He*, Xinkai Zou*, Xuan Jiang, Xingyuan Ding, Ao Qu, Juncheng Billy Li, Heather Miller

    Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026

    2 citations
  4. CloudAnoAgent: Anomaly Detection for Cloud Sites via LLM Agent with Neuro-Symbolic Mechanism

    Xinkai Zou, Xuan Jiang, Ruikai Huang, Haoze He, Parv Kapoor, Jinhua Zhao

    COLM 2026 Workshop on AI Measurement, 2026

    6 citations
  5. 1 citation
  6. Evaluating Multi-Modal UAM Impacts at Metro Region Scale: A San Francisco Bay Area Study

    Xin Peng, Xuan Jiang, Vishwanath Bulusu, Emin Burak Onat, Raja Sengupta

    IEEE Transactions on Intelligent Transportation Systems, 2026

  7. Sparse Matrix in Large Language Model Fine-Tuning

    Haoze He*, Juncheng Billy Li*, Xuan Jiang, Heather Miller

    International Conference on Learning Representations (ICLR), 2025

    38 citations
  8. 15 citations
  9. 6 citations
  10. Large Scale Multi-GPU Based Parallel Traffic Simulation for Accelerated Traffic Assignment and Propagation

    Xuan Jiang, Raja Sengupta, James Demmel, Samuel Williams

    Transportation Research Part C: Emerging Technologies, 2024First author

    36 citations
  11. A Simulation-Optimization Framework for Developing Wind-Resilient AAM Networks

    Emin Burak Onat, Shangqing Cao, Raiyan Rizwan, Xuan Jiang, Mark Hansen, Raja Sengupta, Anjan Chakrabarty

    International Conference on Research in Air Transportation (ICRAT), 2024Best paper

    5 citations
  12. Devastator: A Scalable Parallel Discrete Event Simulation Framework for Modern C++

    John Bachan, Jianlan Ye, Xuan Jiang, Tan Nguyen, Mahesh Natarajan, Maximilian Bremer, Cy Chan

    ACM SIGSIM Conference on Principles of Advanced Discrete Simulation (PADS), 2024

    6 citations
  13. Designing a Time-Driven Simulation Framework for Large-Scale Traffic Networks

    Xuan Jiang

    ACM SIGSIM Conference on Principles of Advanced Discrete Simulation (PADS), 2024Sole author

    10 citations
  14. Simulating Integration of Urban Air Mobility into Existing Transportation Systems: Survey

    Xuan Jiang, Yuhan Tang, Junzhe Cao, Vishwanath Bulusu, Hao Yang, Xin Peng, Yunhan Zheng, Jinhua Zhao, Raja Sengupta

    Journal of Air Transportation, 2024First author

    34 citations
  15. Fleet Size and Spill for UAM Operation under Uncertain Demand

    Shangqing Cao, Xuan Jiang, Emin Burak Onat, Bo Zou, Mark Hansen, Raja Sengupta, Anjan Chakrabarty

    International Conference on Research in Air Transportation (ICRAT), 2024

    17 citations
  16. 9 citations
  17. Entropy-Based Dynamic Programming for Efficient Vehicle Parking

    Jean-Luc Lupien, Abdullah Alhadlaq, Yuhan Tang, Yan Wu, Jiayu Joyce Chen, Yutan Long, Xuan Jiang

    IEEE International Conference on Intelligent Transportation Systems (ITSC), 2024

    2 citations
  18. Integrating the Traffic Science with Representation Learning for City-Wide Network Congestion Prediction

    Wenqing Zheng, Hao Frank Yang, Jiarui Cai, Peihao Wang, Xuan Jiang, Simon Shaolei Du, Yinhai Wang, Zhangyang Wang

    Information Fusion, 2023

    38 citations
  19. An Automated Machine Learning (AutoML) Method for Driving Distraction Detection Based on Lane-Keeping Performance

    Chen Chai, Juanwu Lu, Xuan Jiang, Xiupeng Shi, Zeng Zeng

    INFORMS Annual Meeting, 2023

    22 citations
  20. A Metrics-based Method for Evaluating Corridors for Urban Air Mobility Operations

    Xuan Jiang, Xin Peng, Vishwanath Bulusu, Cristian Poliziani, Gano Chatterji, Raja Sengupta

    IEEE International Smart Cities Conference (ISC2), 2022First author

    12 citations
  21. Causality and Advanced Models in Trip Mode Prediction: Interest in Choosing Swissmetro

    Huy Pham, Xuan Jiang, Cong Zhang

    International Conference on Transportation and Development (ICTD), 2022

    7 citations
  22. Quantifying the Resilience of the US Domestic Aviation Network During the COVID-19 Pandemic

    Aleksandar Bauranov, Steven Parks, Xuan Jiang, Jasenka Rakas, Marta C. González

    Frontiers in Built Environment, 2021

    45 citations
  23. Making Sense of Electrical Vehicle Discussions Using Sentiment Analysis on Closely Related News and User Comments

    Xuan Jiang, Josh Everts

    International Conference on Transportation and Development (ICTD), 2021First author

    17 citations
  24. Chemical and Rheology Evaluation on the Field Short-Term Aging of High Content Polymer Modified Asphalt

    Weidong Huang, Chuanqi Yan, Xuan Jiang

    Transportation Research Board Annual Meeting, 2019

    7 citations

An asterisk indicates equal contribution.

Academic background

Background, honors, and service.

Education

  • 2016 to 2020Tongji UniversityB.S.
  • 2020 to 2024UC BerkeleyM.S. (2020 to 2021) and Ph.D. (2021 to 2024). Specialized in machine learning and HPC, advised by James Demmel, Raja Sengupta, and Alexandre Bayen.

Experience

  • 2026 to presentAmazon Web Services (AWS), Annapurna LabsSr. HPC/AI Engineer. I work on GPU communication performance on large GPU clusters and optimize kernels for mixture-of-experts expert parallelism, and with NVIDIA's NCCL team, I co-lead the work to run NCCL EP on EFA. I co-led DeepEP v2 on EFA and led the NVSHMEM work for DeepEP v1 with NVIDIA's NVSHMEM team, where the changes shipped in NVSHMEM 3.7.0.
  • 2024 to 2026GoogleSoftware Engineer. I led AI-assisted workflows for large-scale code changes and AI-powered integration test setups. I built multi-agent frameworks for agentic code generation and launched alert-based investigations in Google Cloud Assist.
  • 2023 to presentMIT Research Affiliate. With Jinhua Zhao and Haris Koutsopoulos, I work on reinforcement learning to optimize complex systems using HPC-based simulators and on LLM agents for cloud anomaly detection. Since 2026, I have also built self-evolving agents with the reef team and co-authored CORAL (COLM 2026).
  • 2022 to 2025UC Berkeley Graduate Student Researcher at the Aviation Innovation Research (AIR) Lab and the Cal Unmanned Lab (2022 to 2024). Researched HPC with James Demmel and Raja Sengupta using GPU-based multimodal time-driven simulation (2022 to 2025).
  • 2022 to 2023University of Washington STAR LabResearch Assistant. Worked on TinT for city-wide congestion prediction.
  • 2021 to 2023Lawrence Berkeley National LaboratoryGraduate Student Researcher.
  • 2019 to 2020Shanghai Huimice Information Technology Co., Ltd.AI Scientist.

Awards

  • 2024ICRAT Best Paper Award, Trajectories and Networks track.
  • 2023ASCE ICTD AI in Transportation Committee Outstanding Session Organizer Award.
  • 2022NSF AI Workshop Phase II Travel Award.
  • 2021Joseph M. Sussman Best Paper Prize.
  • 2020Gold Award, 6th China International "Internet+" College Students' Innovation and Entrepreneurship Competition.
  • 2019Outstanding Student of Shanghai.
  • 2018Shanghai Municipal Innovation and Entrepreneurship Project Award.
  • 2017Outstanding Student Leader, Tongji University.

Teaching

  • Fall 2022UC BerkeleyGraduate Student Instructor for Control and Information Management.
  • 2023 to 2024UC BerkeleyTeaching Assistant for Emerging Technologies for Public Health.

Academic service

  • 2024Frontiers in Built EnvironmentTopic Editor for the research topic "Advanced Technologies for Aviation Operations and Passenger Experience".
  • 2023 to 2024ASCEMember of the Artificial Intelligence in Transportation Committee and vice chair of its Student Subcommittee.
  • 2023 to 2024ASCEMember of the Active Transportation Committee.

Reviewing

Journal reviewing

  • IEEE Transactions on Intelligent Transportation Systems
  • IEEE Transactions on Industrial Informatics
  • Scientific Reports
  • Frontiers in Psychology

Conference reviewing

  • ICLR 2026
  • ICML 2026
  • CVPR 2026
  • ECCV 2026
  • COLM 2026
  • NeurIPS 2025
  • AAAI 2025
  • ITSC 2025
  • TRB Annual Meeting 2025

Mentoring

Students I have mentored.

Earlier work

Helping build an AI science channel.

Before moving to the US, I helped Zihao (同济子豪兄) build his AI popular-science channel on Bilibili. The channel now has more than 350,000 followers.

Keep in touch

Subscribe, thenlet's be friends.

My weekly newsletter is free. If you work on AI infrastructure, self‑evolving agents or large-scale simulation, you can also email me at xuanj@mit.edu to say hi.