haotian-liu/LLaVA ? reverse-engineered prompt
Reverse engineered prompt
Build me a working version of LLaVA, a chat assistant that can look at an image and answer questions about it in plain English.
I want a simple experience where I can upload a photo or point it at an image file, then ask things like what is in the picture, read text from it, compare objects, or have a back and forth conversation about what it sees. If the project includes a demo or web interface, make that feel polished and easy to use. If there are scripts for running inference or trying different models, wire those up so they actually work from the repo.
Please keep it practical and make sure the main flow runs end to end with sensible defaults. If you need current setup details or model instructions, look up the latest docs online and use the repo’s README and examples as the source of truth.
Are you gonna build this?
make sure you review the code using coderabbit