I2E: From Image Pixels to Actionable Interactive Environments for Text-Guided Image Editing

We propose I2E, a novel "Decompose-then-Action" paradigm that transforms unstructured images into discrete, manipulable object layers and employs a physics-aware Vision-Language-Action Agent to parse complex instructions into atomic actions via Chain-of-Thought reasoning, significantly outperforming state-of-the-art methods on compositional editing tasks.
🌟This work is a tiny step toward my vision of unified intelligence — guiding agents to reason about inter-layer relationships and manipulate objects within a structured visual environment under real-world physical constraints.
BibTeX
@article{yu2026i2e,
title = {I2E: From Image Pixels to Actionable Interactive Environments for Text-Guided Image Editing},
author = {Yu, Jinghan^{1} and Xiao, Junhao^{1} and Zhu, Chenyu^1 and Li, Jiaming^1 and Li, Jia^1 and Deng, HanMing^1 and Wang, Xirui^1 and Jia, Guoli^2 and Li, Jianjun^1 and Ma, Zhiyuan^{1\dag} and Bai, Xiang^1 and Zhou, Bowen^{2,3}},
journal = {ACL 2026 main, arXiv:2601.03741},
url = {https://arxiv.org/abs/2601.03741}
}