chenyu zhuall notes

Welcome to My Research Blog

Notes on unified multimodal models, generative intelligence, and the path toward systems that can see, imagine, reason, and create.

Welcome to My Research Blog

General intelligence demands more than reasoning alone. It also requires the ability to see, imagine, and create. Language, perception, and generation are not isolated capabilities; they are deeply connected parts of a unified mind.

This blog is where I will document ideas, experiments, and open questions from that perspective.

Why Unified Multimodal Models

I am interested in models that do more than translate between isolated modalities. A unified model should be able to perceive a scene, reason about its structure, imagine what could happen next, and create or act with those relationships intact.

That goal connects my work on multimodal large language models, generative models, and interactive visual environments. Each direction offers a different view of the same underlying question: how can an intelligent system build and use a coherent model of the world?

Seeing, Imagining, and Creating

Perception gives a model evidence about the world. Imagination lets it explore possibilities beyond what is immediately visible. Creation turns those possibilities into images, environments, or actions that people can inspect and refine.

My research explores the interaction between these abilities, including structured image manipulation and diffusion post-training. You can find the current projects on the publications page.

Research directions in generative intelligence
Research directions in generative intelligence

What I Will Share Here

Future notes will cover paper reading, experimental lessons, implementation details, and unfinished research ideas. Some posts will be technical; others will be shorter attempts to clarify a question before it becomes a project.

The aim is not to present every thought as a finished result. It is to make the process visible and to leave a useful trail for future work.

Let Us Talk

I am always open to discussions and collaborations around unified multimodal models, MLLMs, world models, and generative learning. If a note connects with your work, feel free to reach out by email or through GitHub.