Jae Sung (James) Park

Research Scientist @ Ai2

I am a Research Scientist at the Allen Institute for AI (Ai2). I completed my PhD in Computer Science & Engineering from the University of Washington, advised by Ali Farhadi, Yejin Choi, and Ranjay Krishna.

Currently, I am interested in multimodal grounded reasoning: how machines use visual perception to ground concepts in images and videos, and reason about the visual world. In particular, I work on:

  • Developing truly open-source multimodal foundation models from scratch (e.g., Molmo).
  • Exploring better and more efficient grounding representations that align with how humans interpret visual content.
  • Building domain-specialized models for multimodal applications and real-world assistants.

Email: jamesp@allenai.org  /  Google Scholar  /  X  /  Github

profile photo
News
  • [06/2026] 2 papers have been accepted to ECCV 2026, including MolmoPoint!
  • [06/2026] 3 papers at CVPR 2026 including Molmo2 (Best Paper Award Candidate), VideoNet (Highlight).
  • [05/2026] Released MolmoAct2: Open Action Reasoning Models for Robot Control and Real-world Deployment! [paper]
  • [03/2026] Released MolmoPoint: Better Pointing for VLMs with Grounding Tokens! [demo] / [twitter]
  • [01/2026] Joined Ai2 as Research Scientist.
  • [12/2025] Defended my PhD at UW!
  • [12/2025] Released Molmo2: Open Image and Video Models with Pointing and Tracking! [demo] / [tracking demo]
  • [06/2025] 2 papers, Molmo and Synthetic Visual Genome at CVPR 2025.
  • [09/2024] Released Molmo, an open state-of-the-art multimodal AI model [demo] [code]
Research
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
Christopher Clark, Yue Yang, Jae Sung Park, Zixian Ma, Jieyu Zhang, Rohun Tripathi, Mohammadreza Salehi, Sangho Lee, Taira Anderson, Winson Han, Ranjay Krishna Core contributor
ECCV, 2026
blog / arXiv / model / data / code / demo / twitter
Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos
Ziqi Gao, Jieyu Zhang, Wisdom Oluchi Ikezogwo, Jae Sung Park, Tario G. You, Daniel Ogbu, Chenhao Zheng, Weikai Huang, Yinuo Yang, Winson Han, Quan Kong, Rajat Saini, Ranjay Krishna
ECCV, 2026
arXiv
MolmoAct2: Action Reasoning Models for Real-world Deployment
Haoquan Fang, Jiafei Duan, Donovan Clay, Sam Wang, Shuo Liu, Weikai Huang, Xiang Fan, Wei-Chuan Tsai ... Jae Sung Park ... Ali Farhadi, Dieter Fox, Ranjay Krishna
arXiv, 2026
blog / arXiv / code
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
Chris Clark*, Jieyu Zhang*, Zixian Ma*, Jae Sung Park*, Mohammdreza Salehi, Rohun Tripathi, Sangho Lee ..., Ali Farhadi, Ranjay Krishna
* Co-first author Core contributor
Led development of tracking capability.
CVPR, 2026 Best Paper Award Candidate
blog / paper / model / data / code /
VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition
Tanush Yadav, Mohammadreza Salehi, Jae Sung Park, Vivek Ramanujan, Hannaneh Hajishirzi, Yejin Choi, Ali Farhadi, Rohun Tripathi, Ranjay Krishna
CVPR, 2026 (Highlight)
arXiv / website / dataset / code
Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
Weikai Huang, Jieyu Zhang, Taoyang Jia, Chenhao Zheng, Ziqi Gao, Jae Sung Park, Winson Han, Ranjay Krishna
CVPR, 2026
arXiv
Synthetic Visual Genome.
Jae Sung Park, Zixian Ma, Linjie Li, Chenhao Zheng, Cheng-Yu Hsieh, Ximing Lu, Khyathi Chandu, Quan Kong, Norimasa Kobori, Ali Farhadi, Yejin Choi, Ranjay Krishna
CVPR, 2025
arXiv / website / dataset / model / code
Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models.
Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi ... Ranjay Krishna, Luca Weihs, Noah A Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, Aniruddha Kembhavi
CVPR, 2025 Best Paper Award Candidate
arXiv / demo / dataset / code
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness.
Khyathi Raghavi Chandu, Linjie Li, Anas Awadalla, Ximing Lu, Jae Sung Park, Jack Hessel, Lijuan Wang, Yejin Choi
ICLR, 2025
arXiv
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Caption
Anas Awadalla, Le Xue, Manli Shu, An Yan, Jun Wang, Senthil Purushwalkam, Sheng Shen, Hannah Lee, Oscar Lo, Jae Sung Park, Etash Guha, Silvio Savarese, Ludwig Schmidt, Yejin Choi, Caiming Xiong, Ran Xu
arxiv, 2024
arXiv / dataset
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
Mohammadreza Salehi, Jae Sung Park, Tanush Yadav, Aditya Kusupati, Ranjay Krishna, Yejin Choi, Hannaneh Hajishirzi, Ali Farhad
Neurips Dataset & Benchmarks, 2024
arXiv / website
Superposed Decoding: Multiple Generations from a Single Autoregressive Inference Pass
Ethan Shen, Alan Fan, Sarah Pratt, Jae Sung Park, Matthew Wallingford, Sham Kakade, Ari Holtzman, Ranjay Krishna, Ali Farhadi, Aditya Kusupati.
Neurips, 2024
arXiv
Localized Symbolic Knowledge Distillation for Visual Commonsense Models
Jae Sung Park, Jack Hessel, Khyathi Chandu, Paul Pu Liang, Ximing Lu, Peter West, Youngjae Yu, Qiuyuan Huang, Jianfeng Gao, Ali Farhadi, Yejin Choi
Neurips, 2023
arXiv
Multimodal knowledge alignment with reinforcement learning
Youngjae Yu, Jiwan Chung, Heeseung Yun, Jack Hessel, Jae Sung Park, Ximing Lu, Prithviraj Ammanabrolu, Rowan Zellers, Ronan Le Bras, Gunhee Kim, Yejin Choi
CVPR, 2023
arXiv
Exposing the limits of video-text models through contrast sets
Jae Sung Park, Sheng Shen, Ali Farhadi, Trevor Darrell, Yejin Choi, Anna Rohrbach
NAACL (short), 2022
arXiv / code
Merlot: Multimodal neural script knowledge models
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, Yejin Choi
Neurips, 2021
arXiv
LLC: Accurate, multi-purpose learnt low-dimensional binary codes
Aditya Kusupati, Matthew Wallingford, Vivek Ramanujan, Raghav Somani, Jae Sung Park, Krishna Pillutla, Prateek Jain, Sham Kakade, Ali Farhadi
Neurips, 2021
arXiv
Natural language rationales with full-stack visual reasoning: From pixels to semantic frames to commonsense graphs
Ana Marasović, Chandra Bhagavatula, Jae Sung Park, Ronan Le Bras, Noah A Smith, Yejin Choi
Findings of EMNLP, 2020
arXiv
VisualCOMET: Reasoning about the Dynamic Context of a Still Image
Jae Sung Park, Chandra Bhagavatula, Roozbeh Mottaghi, Ali Farhadi, Yejin Choi
ECCV, 2020 (Spotlight)
project page / arXiv / code
Identity Aware Multi-Sentence Video Description
Jae Sung Park, Trevor Darrell, Anna Rohrbach
ECCV, 2020
project page / arXiv
Adversarial Inference for Multi-Sentence Video Description
Jae Sung Park, Marcus Rohrbach, Trevor Darrell, Anna Rohrbach
CVPR, 2019 (Oral)
arxiv / code

Service
Teaching