> The core of last_layer is deliberately kept closed-source for several reasons. Foremost among these is the concern over reverse engineering. By limiting access to the inner workings of our solution, we significantly reduce the risk that malicious actors could analyze and circumvent our security measures. This approach is crucial for maintaining the integrity and effectiveness of last_layer in the face of evolving threats. Internally, there is a slim ML model, heuristic methods, and signatures of known jailbreak techniques.
Otherwise I'm left evaluating it through wasting my time playing whac-a-mole with it, which won't give me the confidence I need because I can't be sure an attacker won't guess a strategy that I didn't think of myself.
This doesn't even include details of the evals they are using! It's impossible to evaluate whether what they've built is effective or not.
I'm also not keen on running a compiled .so file released by a group with no information on even who the authors are.
Like a way to block certain google searches you don't agree with.
https://simonwillison.net/2024/Mar/5/prompt-injection-jailbr...
Prompt injection is a security
issue. It’s about preventing
attackers from emailing you and
tricking your personal digital
assistant into sending them your
password reset emails.
No matter how you feel about “safety
filters” on models, if you ever want
a trustworthy digital assistant you
should care about finding robust
solutions for prompt injection.
To be fair, this library attempts to solve both at once.Prompt injection means that even running an LLM against your own private notes to answer questions about them could be unsafe, provided there are any vectors (like Markdown image support) that might be used for exfiltration.
My current recommendation for dealing with prompt injection is to keep it in mind and limit the blast radius if something goes wrong: https://simonwillison.net/2023/Dec/20/mitigate-prompt-inject...
Scope the information that the language model has access to to a subset of the information that the person interfacing with the language model has access to. Prompt injection doesn't matter at that point, because the person will only be able to "leak" information they have permission to access anyways.
More on exfiltration: https://simonwillison.net/search/?q=exfiltration
Edit: related link: https://python.langchain.com/docs/security
Prompt injection is the security flaw that exists because doing that - treating instructions and data as separate things in the context in the LLM - is WAY harder than you might expect.
My previous writing about this: https://simonwillison.net/series/prompt-injection/
Then we should improve the tooling around this to make it way easier, rather than hoping security by obscurity will work this time.