Putting together all I know about tokenminning. Point your agent at it.
1. Prompt hygiene (schema over prose, make no mistake)
2. Cache whenever you can
3. Context hygiene (seriously, start a new chat, it doesn't take much)
4. Routing
5. Control max_tokens, be smart about RAG
6. Semantic caching