cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents
Preprint, 2026.
See full list on Google Scholar
* denotes equal contribution.
cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents
Preprint, 2026.
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
ACL, 2024. As seen on: Wired.
Grounding Language Models to Images for Multimodal Inputs and Outputs
ICML, 2023.
VQ3D: Learning a 3D-Aware Generative Model on ImageNet
ICCV (oral, best paper finalist), 2023.
Text-to-Image Generation Grounded by Fine-Grained User Attention
WACV, 2021.