Clef: Open-source decision models, and new RL fine-tuning platform(blog.cloudflare.com) |
Clef: Open-source decision models, and new RL fine-tuning platform(blog.cloudflare.com) |
And it's only been a few weeks.
There are many, many of those left around, because AI frontier is moving forward so fast, everyone is racing ahead. Which is why I laugh when people say AI is not transformative and LLMs are a dead end (and my favorite, "what are we going to do with all those GPUs when the bubble pops?"). Even if SOTA LLMs hit a hard capability limit tomorrow and never advanced again, there's a good decade of growth and advancement to be extracted just from all the low-hanging fruits that were left unpicked along the way.
The weights have permissive licensing, but the data and training pipeline are not published to reproduce them from their proprietary Qwen starting points. Weights are not "source."
<sad trombone sound>
Surely someone will soon do what the title of this post makes it seem like cloudfare did. Truly modular open source training and inference logic, along with a totally open corpus and weights, will eventually out-compete the closed ecosystem.
I bet the competition will result in research into how to make these decision models several more orders of magnitude faster and cheaper.
Here's a challenge problem - look at a 1M context window and produce N decisions (different queries) from it in 50-100ms.
At0M: A 60M local Jev at 16 ms latency and 79% accuracy on Typed Decision
https://at0m.pienomial.com/ https://news.ycombinator.com/item?id=49920350
We wanted to stress test the system before the V1 release.
Businesses are built on outliers. It doesn't make sense throwing your hard earned insights while paying them money to steal it.
Also, cloudflare https://robindev.substack.com/p/cloudflare-took-down-our-web...
16ms latency. And locally run.
Why go big when you can go small ?
There is no official qwen 3.8 9b
From the model card:
> Clef-Flash is post-trained from Qwen/Qwen3.5-9B. See Clef for the larger variant.
Or are companies/people already building this based on say an arXiv docs? n
---
The pricing is ... hm more expensive but not at the point I won't give it a try due to the embeded vision encoding
At0M: A 60M local Jev at 16 ms latency and 79% accuracy on Typed Decision
https://at0m.pienomial.com/ https://news.ycombinator.com/item?id=49920350
I'd love an privacy first on-device model i could use in iOS.
I built this for my own needs, and thought others might find it useful too.
Perhaps that may be too costly atm
- Jev/TypeSafe: 230 ms median, 254 ms mean
- Jev/OpenRouter: 237 ms median, 267 ms mean
- Clef Flash: 661 ms median, 806 ms mean
What gives?
Instead, to make up for the lack of economic viability of their models, they are forced to release publicly to get marketing to get others to pay based on hype.
Now that we're hitting against the hardware supply limits of global economy, I expect more people to go back and revisit the things left along the way in the mad rush to "just throw more compute at it / make a bigger model" - and thus many more cases like Jev to show up in the next few years.
I've been playing with this for the last year or so. Started with a personal email classifier, also did benchmarks with some public datasets, then created a couple classifiers that could play Doom, and now I've been trying out some other experiments, like a request proxy/router to automatically choose a classifier and fallback to LLM to handle unseen requests
Jev did a great job at creating hype, but also at shaping the concept and space of "decision engine" or "decision model". People were already doing this with LLMs, which is very inefficient for most tasks like that, and the Jev guys figured there was a market there. It seems like they were right, and now there's a rush to flood the space, taking advantage of the hype window
In general, training a general purpose classifier is something lots of people have worked on for a long time. Large Transformer models themselves are typically "generalists" already, so structured generation and constrained decoding have given you the ability to use an LLM as a general classifier for years. It's an incredibly common pattern for working with LLM judges or any sort of branched decision making workflow.
A lot of people who are a bit less familiar with the field saw the hype around Jev and presumed that the reason it was so exciting was that it was a fundamentally new interface for working with an LLM. And that additional excitement drove even more attention to Jev. But fundamentally, TypeSafe's announcement was that they found a particular architecture/training paradigm that resulted in a model for this particular interface that had incredible accuracy, very low latency, and for which they could offer inference at a super low cost.
I've not kept up with the flood of Jev clones that have been released, but I think this is just typical for any new component in deep learning that gets popular. There are an absurd number of open source autoregressive LLMs and fine tunes you can use. The thing that makes one more popular than the other is typically the general performance of the individual model.
But training a model for this purpose, or emulating the procedures described in Jev's papers, isn't something that would be beyond the capabilities of any lab. It's not an entirely alien architecture or approach.
The bigger question for TypeSafe as a company would be if other teams are producing Jev-like models that win on performance or cost. Like I said, I haven't followed the reports super closely, so no idea if that's the case or not.
Latency won't be that good, but could still work similarly. Simply force the structured output of a LLM to the given schema.
Probably also easy to train because we can use stronget LLMs to generate input/output data, or even synthetic data is easy to generate.
It's not really a new technology, it's more like a new use-case.
The hard part is the data and evaluation. Sure, it’s not that hard to build a fast model with good predictive power. But fast at doing what? You probably don’t care about classifying whether a hotdog is a sandwich (which is the Jev demo).
There are a few technical details that can reduce the latency significantly (covered in the post) but the real insight has been from watching the reaction to Jev and seeing that there's enough of a market interest to offer it as a distinct thing. The underlying concept/approach was already there.
The fascinating part to me is that Jev seems like this technique plus post-training to get multiple independent confidence values for each possible answer.
But, for these adhoc models, you need to understand the task more, collect some data and train the model (on CPU, no need for GPU). So Jev-like models are a great way of getting a hosted general decision model, but if you have a very narrow task or set of tasks, you might be better off with some more basic models that you can run on the same server you run other things or even on your laptop
Anyone can copy that and apply to an array of models - stripped down LLMs or already slim/highly performant traditional classification architectures (just wrap inference with an api that inputs/outputs the same structured data).
Jev, I think, would say their advantage is the intelligence of their models and training data including calibration: https://medium.com/code-applied/calibrated-classifiers-makin... (which i still struggle with in the general application... there's no free lunch with these things).
If you have a very narrow use case you can train a BERT based decision model on a laptop an hour if you have good data to train it on. It'll answer faster than the roundtrip to clef/jev and use <1gb memory
Many people seem to have run into the same question and started working out the answer.
It seems insanely obvious at least to me, that JEV is the new hot thing for the AI field since they give you stronger output that isn't... flat out wrong, that alone is impressive.
Getting training data that works well for calibrated classification objectives is difficult.
I hear conflicting opinions (including my own) about how well calibrated each of these are. Jev seems to be the best.
But the jev release made obvious the PMF for these models, and the underlying reality is that calibration really doesn't matter much when you're replacing usecases where people were using damn LM head softmax probabilities before, which are nowhere near calibrated.
So now everyone simply finetunes qwen and makes a compared-to-regular-LLM vastly cheaper decision model. And it works for majority of usecases. People mostly only care about accuracy, not confidence.
We have overcome split brain problems before so this wont be our first
Diffusion transformers are not "easy" but underfunded.
Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware - opens up so many possibilities I'm probably unable to imagine half of them.
E.g. Imagine spellcheck/predictive text (or code autocomplete) where the model is able to process a whole paragraph + surrounding application/system context in between keystrokes. Or an OS being able to reliably guess what you're doing in real-time, in between your UI interactions, and offer actually helpful contextual reactions.
Or imagine finally funding some decent studies into exploring the models as computational artifacts - studying their latent spaces, how they form and how they model reality internally.
Or imagine automated sliding doors that don't suck.
--
[0] - Or anything substantially better than BERT-level models used in Jev or that demo from the company doing inference ASICs, that has a chatbot online that does 14 kilotokens per second.
But also harnesses and more generally new insights on "the control flow problem" could end up squeezing a ton of performance out of small models.
Enough stuff can happen, software use itself might change, and that could really cause anything. "What will we do with all the gpus" might become a question if for a magnitude of tech and reasons leaked-opus-9 runs on a macbook m6 or 7
That's low hanging for you?
This got everyone racing forward and right now there is not enough human attention left in the world to productionize this, or any of the other "side threads". When the race slows down, people will catch up, branch out, and loop back.
[1]: "FPGA-based CNN Acceleration using Pattern-Aware Pruning" https://inria.hal.science/hal-04689673/document
Of course you can also do ranking one-off with a decision model, but this likely less stable, and by doing pairwise ranking you can also relatively quickly do incremental inserts to the list.
Assuming it won't get to full RSI, the current approach will burn out - most likely economically. The race slows down, people branch out, look back, start picking up the "untapped potential"/low-hanging fruits, and you have new S-curves launching in place of the one that just tapered off (hence a fallacy - a stack of S-curves adds up to continuing exponential growth).
In other words: it comes and goes. Hyperconcentrated capital will eventually deconcentrate.