OpenMOSS/MOSS-VL ? reverse-engineered prompt
Reverse engineered prompt
I want this repo turned into a working local demo for MOSS VL so I can actually use it without digging through papers. Make it easy to run on my machine with clear setup steps, then let me do two things. First, I want real time video understanding from either a webcam or an mp4, where I can ask questions while the video is still playing and get answers as it watches. It should keep watching quietly when there is not enough context yet, and update or correct itself if the scene changes. Second, I want an offline mode where I can give it an image or a full video with a prompt and get a normal answer back.
Please use the provided open checkpoints and wire up the repo so the basic inference paths work end to end. A simple local interface is fine, browser or terminal, as long as it feels usable. Include example commands, sensible defaults, and a short README for how to start and test it. Look up current docs online if you need to.
Have a live product UI? Try website reverse