IDEA-Research/GroundingDINO ? reverse-engineered prompt

Reverse engineered prompt

Build me a simple app around Grounding DINO so I can upload an image and type a text prompt like “person, dog, bicycle,” then see the detected objects boxed on the image with labels and confidence. I want it to feel easy to run locally, with clear setup instructions, and if there are good default model weights or a demo already in the repo, use those. Please make sure it works end to end, from loading the model to showing the annotated result, and include a clean way to test it on a few sample images. If anything depends on current docs or the latest model usage, look it up online and wire it up the right way. Also add a small example for how to run it in Docker or with the provided environment files so I can get it working without much fuss.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab