A human-annotated dataset and benchmark of foreground, background, and environment color descriptions for real images, showing that medium-grained color supervision improves vision-language instruction tuning.
Ming-Chang Chiu
I am a Founding Scientist at rekursiv.ai, currently focusing on ensuring our harness can be reliably deployed for discovery in Virtual Cell domain. Previously, I founded Carabin, an AI assurance startup building trust infrastructure for autonomous inter-agent interactions, and in my essay Trust at Machine Speed argued that well-intentioned agents can quietly diverge in what they believe they’ve agreed to, and that both commerce and science need a clearinghouse for the agent era. My academic work has the same instinct for surfacing failures that aggregate metrics hide: subgroup discrepancies in image classifiers that look accurate overall (ICCV 2023), color-contrast effects in machine-assisted skin disease detection (ICASSP 2024), and behavioral inconsistencies in vision-language models. More recently, AIDE (ICLR 2025) explored how multimodal models can agentically recruit domain experts to improve themselves. I was also a contributor to VideoPoet where Josh and Dan also worked on, so it is nice to be with them again!
I was also a Cornell Runway Postdoc and hold a PhD in Computer Science from the University of Southern California, where I was advised by Xuezhe Ma and collaborated closely with Pin-Yu Chen at IBM Research. I have published at venues including ICLR, ICCV, AAAI, and ICASSP, review for ICML, NeurIPS, and ACL, and previously conducted research at Google Research, NVIDIA, and Lawrence Livermore National Lab. I am passionate about ensuring AI systems can verify and collaborate safely at scale.
Outside of work I backpack, snowboard, hunt for vintage finds, and listen to both classical and rock.
Selected Publications
My research sits at the intersection of trustworthy AI, vision-language models, and agentic systems. For a complete list, see my Google Scholar (690+ citations, h-index 6).
A framework in which a vision-language model selects its own weak instances, recruits specialized domain-expert models as tools, and synthesizes their outputs into new training data to improve itself without larger teacher models.
A benchmark of human-annotated color labels for ImageNet and CIFAR that measures how strongly machine recognition depends on color, and shows color-aware augmentation improves robustness.
A multimodal large language model for zero-shot video and audio generation from a wide variety of conditioning signals including images, videos, text, and audio.
An end-to-end framework that probes vision-language models for human-like behavioral biases such as recency bias and authority bias using real stock-price and earnings data.
A dataset annotating skin-lesion images by color contrast, revealing that models perform worse on low-contrast lesions and that the effect compounds across skin tones.
Shows that image classifiers with strong overall accuracy still exhibit large gaps across natural subgroups such as object color, and that standard augmentation does not close them.
A multimodal system that segments, diarizes, and transcribes audio-recorded oncology visits, then uses lexical and acoustic features to detect empathic opportunities and expressed empathy between clinician and patient.
Frames screenplay quality as predicting major film-award nominations from linguistic cues, and shows that narratology-inspired domain features improve over strong baselines on long, sparsely labeled scripts.