Ming-Chang Chiu

I am a Founding Scientist at rekursiv.ai, currently focusing on ensuring our harness can be reliably deployed for discovery in Virtual Cell domain. Previously, I founded Carabin, an AI assurance startup building trust infrastructure for autonomous inter-agent interactions, and in my essay Trust at Machine Speed argued that well-intentioned agents can quietly diverge in what they believe they’ve agreed to, and that both commerce and science need a clearinghouse for the agent era. My academic work has the same instinct for surfacing failures that aggregate metrics hide: subgroup discrepancies in image classifiers that look accurate overall (ICCV 2023), color-contrast effects in machine-assisted skin disease detection (ICASSP 2024), and behavioral inconsistencies in vision-language models. More recently, AIDE (ICLR 2025) explored how multimodal models can agentically recruit domain experts to improve themselves. I was also a contributor to VideoPoet where Josh and Dan also worked on, so it is nice to be with them again!

I was also a Cornell Runway Postdoc and hold a PhD in Computer Science from the University of Southern California, where I was advised by Xuezhe Ma and collaborated closely with Pin-Yu Chen at IBM Research. I have published at venues including ICLR, ICCV, AAAI, and ICASSP, review for ICML, NeurIPS, and ACL, and previously conducted research at Google Research, NVIDIA, and Lawrence Livermore National Lab. I am passionate about ensuring AI systems can verify and collaborate safely at scale.

Outside of work I backpack, snowboard, hunt for vintage finds, and listen to both classical and rock.

Ming-Chang Chiu

Selected Publications

My research sits at the intersection of trustworthy AI, vision-language models, and agentic systems. For a complete list, see my Google Scholar (690+ citations, h-index 6).

AIDE
AIDE: Agentically Improve Visual Language Model with Domain Experts
M.-C. Chiu, F. Liu, K. Sapra, A. Tao, Y. Jacoob, X. Ma, Z. Yu, G. Liu
ICLR Workshop on Self-Improving Foundation Models, 2025  ·  4 citations

A framework in which a vision-language model selects its own weak instances, recruits specialized domain-expert models as tools, and synthesizes their outputs into new training data to improve itself without larger teacher models.

ColorSense
ColorSense: A Study on Color Vision in Machine Visual Recognition
M.-C. Chiu, Y. Wang, D. E. G. Kim, P.-Y. Chen, X. Ma
IEEE SaTML, 2025  ·  7 citations

A benchmark of human-annotated color labels for ImageNet and CIFAR that measures how strongly machine recognition depends on color, and shows color-aware augmentation improves robustness.

VideoPoet
VideoPoet: A Large Language Model for Zero-Shot Video Generation
D. Kondratyuk, L. Yu, X. Gu, J. Lezama, J. Huang, G. Schindler, R. Hornung, V. Birodkar, J. Yan, M.-C. Chiu, et al.
ICML, 2024  ·  614 citations / Best Paper

A multimodal large language model for zero-shot video and audio generation from a wide variety of conditioning signals including images, videos, text, and audio.

Automated Empathy Detection
Automated Empathy Detection for Oncology Encounters
Z. Chen, J. Gibson, M.-C. Chiu, Q. Hu, T. K. Knight, D. Meeker, J. A. Tulsky, K. I. Pollak, S. Narayanan
IEEE ICHI, 2020  ·  20 citations

A multimodal system that segments, diarizes, and transcribes audio-recorded oncology visits, then uses lexical and acoustic features to detect empathic opportunities and expressed empathy between clinician and patient.

Screenplay Quality Assessment
Screenplay Quality Assessment: Can We Predict Who Gets Nominated?
M.-C. Chiu, T. Feng, X. Ren, S. Narayanan
ACL Workshop on Narrative Understanding, Storylines, and Events, 2020  ·  4 citations

Frames screenplay quality as predicting major film-award nominations from linguistic cues, and shows that narratology-inspired domain features improve over strong baselines on long, sparsely labeled scripts.