piskvorky/gensim ? reverse-engineered prompt

Reverse engineered prompt

Build me a Python library for topic modelling and document similarity that can handle big text collections without loading everything into memory at once.

I want it to let me feed in a stream of documents, train common models like LDA, LSA, and word embeddings, then query for similar documents or words in a simple way. Make the public API feel easy to use, with clear examples and a quick start that a beginner can follow. Include enough tests and docs so I can understand how to train on my own corpus and reuse the trained models later.

Please keep it efficient, work well with large datasets, and make sure it can run on modern Python versions. If you need to check current docs or best practices online, go ahead and do that.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab