Typesafe AI raises $870M at $7.5B(typesafe.ai) |
Typesafe AI raises $870M at $7.5B(typesafe.ai) |
We released an Apache-2.0, open-weight 4B decision model that scores above Jev 1.13 on JevBench's composite score (72.5 vs 71.5) and is currently the top open model there: https://benchmarkheaven.com/jev-models . Newer models coming even larger than beat Jev in intelligence as well.
- Same contract as Jev: state + typed questions in, calibrated probabilities out, one forward pass, no generated tokens. - Your data never leaves your environment, and there's no per-call fee.
Weights, card and run instructions: https://huggingface.co/h2oai/h2o-lightning-4b
EDIT: As I wrote this Microsoft just released their own Decision-1 model [3].
[1] https://developers.openai.com/api/docs/guides/decisions
[2] https://unsloth.ai/docs/basics/train-your-own-decision-model...
[3] https://commandline.microsoft.com/microsoft-decision-1-model...
There's the potential for an inverse LLM play. In contrast with LLMs, all that seems to matter is the model, and the products the labs build around the models are all really samey and boring; the same left panel list of agents, main view agent conversation, right hand extra context, and we're now in the era of everyone creating the same cutesey furry friend on top of all this tech.
But even if others surpass them and make better solutions, the fact that nobody was able to see it before typesafe is a testament of what they might be able to come up with next.
I sound like a fanboy but I swear I 'm unaffiliated with typesafe. I was building my own version of this way before they announced JEV (mine was ALE and it was mentioned here on HN for a bit), in use for VR gaming (so one can give commands to NPCs with voice and supports multiple commands in sequence in a single pass), but I missed the "killer usecase" of being a new primitive, like everyone else.
TLDR: There's a lot of value in thinking ahead and seeing the future. The clones are nice and exciting but they give me "I could have built this first, yes but you didn't" vibes. I hope they manage to keep it up and push the space forward again
https://seldon-ai.com/blog/fronter-llms-are-semantic-interpr...
Precisely this. Should be top comment.
Also, the model moat is understated as training data for these purposes also accrues to the winner, which due to the first mover advantage as well as the distribution advantage you speak of, is typesafe. In contrast to relatively open coding data. Openai anthropic also have that, but like you say its a different business.
I give it to them for creating the hype (good marketing), and for making a useful classifier. Not sure what they would scale rapidly though.
That’s a different set of properties in terms of size, cost, and latency. Jev (apparently, not like I’ve seen its insides) is extremely cheap, extremely fast, doesn’t cost any output tokens as it speaks the output natively, can’t get the output wrong because it speaks the format natively, and (presumably based on the docs), the context is separate from the question, meaning it should be immune (or at least highly resistant) to prompt injection attacks.
It’s not just about the accuracy of the result, it’s a collection of all the properties that make Jev interesting.
Jev took years to develop, I strongly doubt that a copycat that was put together within days after Jev’s release will be able to match it on a sun of its properties. Even if fine tuned LLMs can outperform it on raw accuracy.
But why? If the simplest way to achieve Jev's capabilities (accuracy, cost, latency) is by fine tuning a small model, what makes you think that this isn't exactly what Typesafe did?
And even if they did something different - what makes you think it was a good idea in the first place, given how easy their results were replicated without any "secret sauce"?
no, you can't, and it's unclear why you would think this.
*if you have a sufficiently sized and quality dataset for the specific classifications you're targeting
Here's a hint: confidence is not generated by a model.
Have you heard that from a different source than OpenAI? From what I'd heard other models haven't gotten close, and the open source ones are like running gemma4 E2B against Opus 5.5- sure, the API calls go in and are returned the same but the quality isn't close.
But, for OpenAI this is not a primary business, for open source models as well, so they will not be chasing the market and customers to buy their product and promise them to maintain it.
TypeSafe will do all this, they will try to understand your use cases and then solve your pain point, while others are providing raw material.
such an arrangement can end up beneficial to the VC firm
To run the SDK examples below, use these OpenAI SDK versions or later: Python 3.26.0,
I thought Pythin 3.15.0 just came out, 3.26.0 must be really far off?
Jev is 42$/B but OpenAI is 100$/B token.
By all means, become an A16Z LP.
They may well be a good team to throw money behind if you are hoping to bet on a new AI lab.
We've yet to see whether this is true, or is it just manufactured demand. There are dozens of Jev demos, but pretty much all of them are either cool but useless, or simply fake (i.e. harness doing 99% of the work).
If anyone else came up with the same concept on a Reddit thread (they have) it no-one would care without those characteristics even if you are "first".
Rebranding, execution, marketing, ex-<big_name_company> and mostly importantly, hype is what gets the investors scrambling into throwing money at you.
I have no faith in the technique if it cannot do the basics (i.e. not real probabilities, the confidence for coin flip outcomes)
tried it a couple of days ago here: https://jevplayground.com
the "not real probability" disclaimer only appears after you get a result
https://www.youtube.com/watch?v=xNgQtzEl4lY
Jev is used as an example of a successful marketing launch where they worked with many X "creators" prior to its release, so that all the creators would repost to put it to the top of everyone's feed. Then, over the following days they'd repost so it maintained momentum.
See: doomers.ai, clickstrike, growth matrix, etc. They use coordinated engagement, paid influencer networks, customized messaging, etc.
Jev isn't a terrible product, but it's way overhyped.
I see a lot of people parroting the quick open source alternatives as being better on the benchmarks, but it's such a new category that I'm not convinced we have solid benchmarks.
I'm hoping a company releases an internal eval benchmark for these options. I'm sure some of the open source ones are solid in some cases, but would love to see more reliable data.
If being an “AI Researcher” is a ticket to multimillion dollar salary, AI training talent cannot be contained to a handful of companies. It’ll become more common and diffuse. The old advice of not fine tuning, because it’s hard, goes out the window as that knowledge diffuses through the industry.
A similar thing is happening in search. For a long time labs have trained tailored embedding models. And now companies like SID training their own agentic models that are smaller and faster at search than GPT-5.
I am integrating it into the product I am building and to me it doesn't seem like there is much need to go with a SaaS for this since the requirements are so light. I just can run it in Cloud Run and get all of the scale I'll ever need, and I get to tell my customers their data never leaves my environment.
I actually resized my browser thinking maybe something weird was going on with flex-wrap or overflow or whatever it is.
A lot of people are shouting about how Jev hasn't actually differentiated itself, but I question how much folks are actually experimenting with what's out there before coming up with an opinion.
For us, it's cleae that OpenAI rushed this out to meet the hype in the market right now without having a product that actually meets the bar Jev has set.
I did. Originally I had a project that I had been wanting to do and thought to use a decision model for it. Jev, OpenAI, etc. are all within percentage points of each other.
Then I used traditional ML and found a small classifier (gemma 4) with traditional embeddings worked 2x as well.
Jev is the general purpose ML pipeline for when you want average results. Nearly every application has a "better" option available with a small amount of work.
Unless china takes leadership in frontier space the picture is next :
1. cheap workhorses for classification, routing, other scenarios : Jev 2. coding agents with less erros : Anthropic/Openai, etc. 3. Science /Legal/Medical : A mixture of Jev+Anthropic scenarios
There's your problem. The single biggest thing every Canadian VC is trying to figure out is "why are these people asking us for money when if they were any good they'd be in the US" so by simply asking them you're already signalling something bad. A lot of their enthusiasm for process is based on this suspicion and also that the entire industry is just a way for various professional services to extract most of the investment money, since that's the game they're so used to playing with the government.
There are some Canadian VCs earnestly trying to improve but they are overwhelmingly hilariously conservative and focused on unimportant signals over reality. This is one (but not all) of the major factors that drive basically every remotely ambitious Canadian company to run a corp in Delaware and go for funding from the US. The tax situation is the other major contributor.
But that does nothing to make up for the terrible investment community. Getting started here requires already being started.
When I briefly worked for a Toronto startup, it was like all of them went to the same private boy's schools together as kids. It was a status club.
I jumped ship to an American startup and made almost double the money dealt with 0% of the bullshit and they were bought by Google the next year.
Also the situation isn't static. Investors know that the act of writing them a $870M check itself increases the chance that they'll be one of the winners, because that will attract more talent, customers, and funding to the company in a self-reinforcing cycle. And investors know that other investors know that, and that someone is going to write them that $870M check, so to some extent they're forced to think of the company as having already been successful at the fundraising and already having that momentum boost.
Only a small number of investors in the world can play the game at this level, because you have to smart enough to be right (often enough), and you have to be established enough to see the deals (be on every CEO's short list - because CEOs are only going to seriously pitch 5-10 VCs on a hot deal, if that). Otherwise you can't pull it off. Martin Casado and his team are among the few that can and I think their results reflect that.
[1] not really but they did some innovative things and popularized a concept
no, it was not
> was duplicated within a couple of days
was it already available or did it become available in a couple of days? it cant be both (neither is true, actually)
I think the key differentiator was that a team found a whitespace in what ChatGPT was doing, main comes from the same pedigree and team is as conscious of marketing as their product. SF VCs love these out of the box challengers, and people are claiming to replicate doesn't seem to matter.
The amount raised feels surprising but again entire SF/US AI scene is primarily "add moar layers and GPU" one trick ponies at this point.
Yes, anyone can wrap a decisions API around an LLM, but so what? If you want to compete then you need to compete on price, and it's not clear if OpenAI and/or Anthropic are able or willing to do that without building a custom architecture, and even then if a race to the bottom on pricing is what they really want to pursue.
I'm not sure if OpenAI have announced pricing for their Decisions API, but they have said it's based on Luna which costs $0.10/M input, not even remotely competitive with Jev's $0.04/M input, which I'd expect has some headroom built into it.
Assuming that the architecture behind Jev is not just an LLM, and gives them some inherent efficiency/cost and speed advantage, then the question is whether OpenAI and Anthropic really want to duplicate this and have a race to the bottom on pricing for what may be a large part of the business automation market they are addressing. Is that what they want as their IPO pitch - we're selling potatoes, and think can grow them cheaper than Typesafe ?
Acquired in the vc sense… not literal exit.
Maybe in some cases. But counterexample, courtesy of The Information:
"It took just 15 minutes for Blue Owl executives to agree to invest up to $10 billion in future projects alongside real estate firm Primary Digital Infrastructure during their first in-person meeting two years ago, said Primary chief investment officer Bill Stein."
https://www.theinformation.com/articles/blue-owl-eyes-new-de...
AI seems to make some people lose their damned minds.
When you are in the middle of a boom cycle, it's the hottest company that has the advantage. Investing in them is a matter of privilege and they get to pick and choose.
Also, Canadian VCs are bottom of the barrel as far as VCs go.
And if you could put the words "Bay Area" or "Stanford" or "San Francisco" next to your name... different story.
The VCs are not buying the idea or the tech, they're investing in the people. And they invest in a formula that has already worked for them before to make big coin. Prop somebody up, let them hire like crazy, and then get them get acquired, and then cash out. They don't care if it fails if they can make it succeed 1/200 times.
Canadian investors want you to have already succeeded before they help you succeed a tiny bit more.
Anyone know when this "have no moat" meme appeared? Even 5 years ago I don't remember seeing it on every post.
I would argue Dropbox did have a moat. It didn't merely store your data. It made it possible to make backup efficiently when bandwidth wasn't all that good.
Reading "no moat" so often is also tied to the fact those companies happen to be getting surreal valuations, at a quite early stage, showing no profit, building a tech that doesn't seem difficult to reproduce.
Plenty has been said about this claim. If you're still falling for this, I feel sorry for you.
If you remove wheels from your car, your car will get a "no speeding ticket" property, and yet there's nothing exciting about it.
> You can’t get that with a fine tuned LLM
Of course you can. All these claims are nothing but marketing.
VC is a hits business. Just one hit pays for 9 that didn't work out.
"SalesRLAgent: A Reinforcement Learning Approach for Real-Time Sales Conversion Prediction and Optimization"
Jev is a general-purpose thing. That is a specific-purpose thing. General-purpose thing is not the same as specific-purpose thing. What makes people think these are the same thing? I don't get it.
Literally misses the point of Jev, which you don't need to fine-tune to get accuracy nor - and no other model has this - some sort of out of sample calibration
Jev is:
- accurate
- general purpose
- fast and cheap
Models we had before Jev had at most 2/3 of above qualities, but none of them were 3/3.
Copypasta:
• A reinforcement learning architecture specifically designed for sales conversation analysis and conversion prediction
• A synthetic data generation pipeline leveraging GPT-4O to create diverse and realistic sales conversations
• Novel state representation techniques using Azure OpenAI embeddings (3072 dimensions) with sales-specific features
• A meta-learning approach enabling the system to express confidence in its predictions based on conversation similarity to training data
• Integration mechanisms providing real-time guidance within existing sales platforms
• Extensive comparative evaluation demonstrating significant performance improvements over LLM-based approaches
Counterpoint: there are a lot of one-hit wonders, and they vastly outnumber the idea-factory people. This is not to minimize those people, a single idea can be very successful (see Zuckerberg), but it doesn't mean your subsequent ideas will also be great (see Zuckerberg)
Here if anyone is interested: https://pantel.is/projects/ai-gaming-companion/
This all sounds intelligent and likely, and yet we can come up with countless counter examples where the first mover is not the big winner, and nobody cares about who did it first. There is typically way more value in nailing the execution of a big idea someone else came up with, rather than "seeing the future".
Do you think though that Luna being a model post-trained for chat produces over-confidence in logprobs?
If the word “calibration” in probability doesn’t mean anything to you, the difference isn’t very apparent, but that doesn’t mean it doesn’t exist.
https://benchmarkheaven.com/jev-models
According to this benchmark, Jev is currently trailing Quyet-1.0-Large and a few other hastily put-together LLM-based decision API-like setups.