Researchers can now reverse-engineer LLM prompts from output text with near-perfect accuracy
Researchers at IIT Bombay and Adobe Research have built an inverse language model that reconstructs the original prompt from an LLM's output with near-perfect accuracy. Their method, called "Previous-Token Prediction," doesn't need access to model weights and works across different models. For companies relying on proprietary system prompts, this could be a serious security risk.
The article Researchers can now reverse-engineer LLM prompts from output text with near-perfect accuracy appeared first on The Decoder.
This newsroom publishes only an opening to its feed and keeps the full text on its own site. The complete article is at the source below.