Jeomon/Windows-Use ? reverse-engineered prompt
Reverse engineered prompt
Build me a Python tool for Windows that lets an AI agent control the desktop through the normal GUI, not by taking screenshots and guessing. I want to be able to give it a task in plain English like open an app, type into it, scroll, switch windows, use keyboard shortcuts, run PowerShell commands, and read what’s on screen through the Windows accessibility tree.
Make it work with a few popular LLM providers, and include a simple terminal command so I can start chatting with the agent from the command line. It should support browser use too, plus basic memory across steps, and voice input and output if that’s easy to add. Please include a clean Python API, a few examples, and tests so it feels solid. If you need to check current docs for any Windows or LLM APIs, go ahead and look them up online.
Are you gonna build this?
make sure you review the code using coderabbit