NVIDIA-AI-IOT/live-vlm-webui ? reverse-engineered prompt

Reverse engineered prompt

Build me a simple web app that lets me open my webcam in the browser and chat with a vision language model in real time.

I want to see the live camera feed, send frames to the model continuously, and get back useful answers about what it sees, like describing the scene, answering questions, and reacting to changes as they happen. It should feel like a clean local tool I can open at https://localhost:8090, with camera permission prompts, a straightforward landing screen, and a chat style interface for prompts and responses.

Please make it work with a backend model setup like Ollama, vLLM, or a cloud API, and include the basics for benchmarking or comparing performance if that is already part of the project. Keep it easy to run on desktop systems and Jetson style hardware if possible, and if you need to check current setup details online while building, go ahead and look them up.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab