Analysis updated 2026-07-27 · repo last pushed 2026-01-13
Find papers and benchmarks that measure whether video AI models can reason through visual puzzles or games.
Discover methods for teaching video models to reason step by step by generating intermediate visual frames.
Locate open-source code repositories for video reasoning techniques to kickstart a product feature.
Survey the academic landscape of video reasoning research before starting a new project in this area.
| gvclab/awesome-reasoning-via-vdm | 0labs-in/vision-link | 1038lab/agnes-ai | |
|---|---|---|---|
| Stars | 4 | 4 | 4 |
| Language | — | TypeScript | Python |
| Last pushed | 2026-01-13 | — | — |
| Maintenance | Quiet | — | — |
| Setup difficulty | easy | moderate | easy |
| Complexity | 1/5 | 3/5 | 2/5 |
| Audience | researcher | developer | vibe coder |
Figures from each repo's GitHub metadata at analysis time.
No setup required, it is a curated list of links in a README file.
This repository is a curated reading list for anyone interested in whether AI can reason through video. It collects recent academic papers and open-source projects that test how well video models can think through problems, rather than simply describing what appears on screen. The collection is split into two parts. The first covers benchmarks and evaluation studies, essentially tests researchers use to measure whether a video model can solve tasks like navigating a maze, playing chess, or working through visual puzzles it has never seen before. The second part covers methods, highlighting emerging techniques for teaching video models to reason step by step, often by generating intermediate visual frames as part of the thinking process. People who would find this useful include AI researchers, graduate students, or product teams exploring video understanding features. For example, if you are building a tool that needs to analyze a how-to video and explain each step, or evaluating whether a model can follow a game of chess from footage alone, this list points you to relevant papers and code repositories to start from. The README is essentially a set of organized links with no narrative explanation, so you will need to follow the links to each paper or project page to learn the details. It reads more like a living bibliography than a standalone guide, which is common for academic "awesome lists" that track a fast-moving research area.
A curated reading list of academic papers and open-source projects about whether AI can reason through video, measuring how well video models think through problems rather than just describing what appears on screen.
Quiet — no commits in 6-12 months (last push 2026-01-13).
No license information is provided in this repository.
Setup difficulty is rated easy, with roughly 5min to a first successful run.
Mainly researcher.
This repo across BitVibe Labs
Verify against the repo before relying on details.