Hetzner is working on LLM Inference(sliplane.io) |
Hetzner is working on LLM Inference(sliplane.io) |
Straight up the opposite, which the name makes abundantly clear, with the option it does reasoning, without it it doesn't...
This allows both the client and server to customize it. I typically use a message that tells it to compact the conversation and use subagents. I find the reasoning gets bloated when it's failed to do whatever task it's doing and often times it either has too little context (subagents) or its context is bloated (compact).
This works fairly well to get it to extend workable life up to ~1M on a local model.