Skip to content

Navigation Menu

Sign in
Sign up

UniX-AI-Lab


GitHub Org Non-Profit License Active



Towards Real-World Perception and Modeling โ€”
building open AI systems that perceive, understand, generate, and act.



๐Ÿงญ Who We Are

UniX-AI-Lab is an open, non-profit research collective working toward:

Towards Real-World Perception and Modeling

We build intelligent systems that can continuously perceive, understand, model, and interact with the real world.

Our research spans four closely connected directions:

Multimodal Perception & Understanding ยท Visual Generation & Editing
Unified World Modeling ยท Agentic Reasoning

We are particularly interested in Multimodal Large Language Models (MLLMs) for video understanding, streaming perception, long-context multimodal reasoning, and real-world interaction.

We believe meaningful progress requires connecting perception and generation: an intelligent system should not only recognize what is happening, but also understand temporal dynamics, predict future states, reason about actions, and model how the world evolves.

Our work combines academic rigor with practical capability. We care deeply about why things work, not only that they do. Every model, dataset, benchmark, and codebase we produce is released openly whenever possible.


๐Ÿ”ฌ Research Pillars

๐Ÿ‘๏ธ Multimodal Perception & Understanding

Building multimodal systems that understand real-world visual and temporal information.

Topics we explore:

  • Multimodal Large Language Models (MLLMs)
  • Image and video understanding
  • Streaming and online video understanding
  • Long-video and long-context reasoning
  • Temporal grounding and event understanding
  • Audio-visual-language alignment
  • Embodied and interactive perception

๐ŸŽจ Visual Generation & Editing

Developing generative models that create, edit, and simulate realistic visual worlds.

Topics we explore:

  • Image, video, and 3D generation
  • Diffusion and flow-based models
  • Instruction-following visual editing
  • Controllable and compositional generation
  • Identity, layout, motion, and style control
  • 3D-consistent generation
  • Reinforcement learning for generative models

๐ŸŒ Unified World Modeling

Unifying perception, generation, and prediction to model how real-world states evolve.

Topics we explore:

  • Unified understanding and generation
  • World models and future-state prediction
  • Any-to-any multimodal architectures
  • Physical, social, and causal reasoning
  • Multimodal tokenization and representation learning
  • Cross-modal grounding and alignment
  • Interactive environment modeling

๐Ÿค– Agentic Reasoning

Building systems that reason, plan, use tools, and act in dynamic environments.

Topics we explore:

  • Multimodal autonomous agents
  • Tool-augmented reasoning
  • Long-horizon planning and decision-making
  • Multi-agent collaboration
  • Reinforcement learning for reasoning
  • Scientific discovery automation
  • Agents grounded in streaming perception

๐Ÿš€ Featured Projects

๐ŸŒ WorldReasonBench

Human-aligned stress testing of video generators as future world-state predictors.

WorldReasonBench reframes video generation evaluation as world-state prediction: given an initial state and an action, can a model generate a future video whose evolution remains physically, socially, logically, and informationally consistent?

๐ŸŽฌ 436 curated cases ยท 4 reasoning dimensions ยท 22 subcategories
๐Ÿค– 11 closed- and open-source generators benchmarked head-to-head
๐Ÿง‘โ€โš–๏ธ WorldRewardBench: ~6K expert preference pairs over 1.4K videos
๐Ÿ“ˆ ScorePR โ†” human ฯ = 0.955

Video Generation World Model Benchmark Reward Model

Stars arXiv HF Paper Dataset

๐ŸŒ Project Page

๐ŸŽฌ StreamOPD

A post-training recipe with spatio-temporal cue gating for streaming video understanding.

StreamOPD fixes a deliberately austere inference path โ€” four recent frames, no memory bank, no retrieval, no reasoning trace โ€” and asks how far post-training alone can go. ST-CueGate then gates on-policy distillation by how much a grounded visual cue actually shifted the teacher's likelihood on the student's own response.

๐ŸŽฏ 84.6 StreamingBench ยท 69.3 OVO-Bench macro โ€” a 4B student past its own 9B teacher ๐ŸŽž๏ธ 4 recent frames at 1 fps ยท memory-free, retrieval-free inference ๐Ÿงช 25K verifiable video QA items with the full construction pipeline released ๐Ÿค— Weights on the Hub, evaluable straight from the repo id

Streaming Video MLLM On-Policy Distillation Post-Training

Stars arXiv HF Paper Model

๐ŸŒ Project Page

๐Ÿ”ฎ More Coming Soon

We are actively developing projects toward real-world perception and modeling.

In the pipeline:

  • Long-video perception and temporal reasoning benchmarks
  • Unified vision-language understanding and generation
  • Instruction-driven video generation and editing
  • Interactive world models for future-state prediction
  • Multimodal agents for real-world reasoning

Watch this space โ€” or better yet, join us.


๐Ÿ’ก Our Philosophy

Open Research โ‰  Slow Research
Non-Profit โ‰  Low Quality
Accessible โ‰  Trivial

We operate by a few core principles:

Principle What it means in practice
Radical openness Code, weights, data, and evaluations โ€” public whenever possible
Depth over breadth We would rather understand one thing deeply than five things superficially
Community first Research should be reproducible, readable, and useful to everyone
No gatekeeping Research direction should be guided by scientific value and community impact
Frontier, not incremental We pursue fundamental contributions rather than marginal benchmark gains
Perception meets modeling Understanding the real world requires connecting observation, prediction, generation, and action

๐Ÿ“„ Publications & Preprints

Publication list coming soon. All work will be posted to arXiv and linked here.

We target top-tier venues including NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, ACL, and EMNLP, and release preprints as soon as they are ready.


๐Ÿ› ๏ธ Tech Stack


๐Ÿค Join Us

UniX-AI-Lab is open to researchers, engineers, and students who share our mission.

Ways to contribute:

  • ๐Ÿ› Issues & PRs โ€” bug fixes, new features, and documentation improvements
  • ๐Ÿ’ฌ Research discussions โ€” open an issue to discuss ideas before implementing
  • ๐Ÿ“ Writing โ€” help us write papers, blog posts, or tutorials
  • ๐Ÿ” Reproduction โ€” replicate prior work and share what you find
  • โญ Star โ€” help other researchers discover our work

We especially welcome:

  • ML researchers at any career stage, including PhD students
  • Engineers who care about open-source research infrastructure
  • Researchers interested in multimodal perception, video understanding, generation, world models, and agents
  • Anyone passionate about responsible and transparent AI development

๐Ÿ“ฌ Contact & Links

๐Ÿ™ GitHub github.com/UniX-AI-Lab
๐Ÿ“ง Email Open an issue โ€” we respond to all of them
๐ŸŒ Website Coming soon

UniX-AI-Lab is an independent, non-profit research collective.
We have no corporate sponsors and no commercial agenda โ€” just curiosity and open science.

If you use our work, please cite it. If you improve it, please share it.



visitors

Pinned Loading

  1. StreamOPD StreamOPD Public

    [Preprint 2026] StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding

    Python 20

Repositories

Loading
Type
Select type
Language
Select language
Sort
Select order
Showing 4 of 4 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading...

Most used topics

Loading...

AltStyle ใซใ‚ˆใฃใฆๅค‰ๆ›ใ•ใ‚ŒใŸใƒšใƒผใ‚ธ (->ใ‚ชใƒชใ‚ธใƒŠใƒซ) /