Muse Spark 1.3(developer.meta.com) |
Muse Spark 1.3(developer.meta.com) |
I have had sos-so results with their local muse 30b model, but the hosted API is very fast and I have been getting good results.
Any company without bad intentions would do the opposite, but no not with Meta. I'm super impressed with their level of evilness on every product.
Only way I see is if it becomes the new SOTA / frontier, does anyone think Meta will surpass Anthropic or OpenAI?
I still can’t get my head around why language models are an existential threat to Meta - they own the platforms people watch adds on?
i like it very much. it is different than all other chinese models distilled from claude.
just ask it to do some front-end work and you will see its not the same UI as all other claude/distils.
also the price is unbeatable, $0.002 input caching. its the same as old dsv4-flash prices.
Given this is Meta, my immediate assumptions that one is cheap because it lets me "be the product". I know I'm rushing to conclusions but there is zero trust here. The brain will do its thing. And the wording here is giving the brains a lot of wiggle room.
I don't see the wiggle room at all.
In other words, it's not that Meta really wants your data and they're willing to pay top dollar for it. It's that companies really don't want Meta to have their data and they're willing to pay top dollar for that.
However, the statement did not include any details, so it is not clear if the open weights variant will be the same that they are hosting now, or some scaled down version.
This is interesting
I feel the same about Grok w/ Elon. I will pay extra to use someone else.
I'm not an Amodei stan, but of all of these people he seems to have the most ethical focus. Again, not everything done perfectly and I have my gripes, but of the leaders of frontier labs, I'll vote with my money.
And, yeah, I wouldn't trust sama to watch my bag while I went to the bathroom.
Anthropic leadership repeatedly presents themselves as uniquely morally qualified to steward agi and decide how humanity should get access to it. Yet they have repeatedly failed basic morality tests.
Pirating books for financial gain. The newer Sony/Warner music case shows this is pattern behavior.
Aggressively scraping other people's works, despite the authors' requests not to do so.
Then applying massive usage restrictions on their own work.
And probably the most disqualifying is backing away from their own hard AI safety commitments.
They want to position AI as an insurmountable threat in order to regulate away any future competitors. They’re trying to speedrun regulatory capture.
From what I know, the "books3" dataset was normalised in the LLM and research ecosystem, where collected datasets were seen as valid to train on and/or fair use. I'm not sure any of the major frontier companies are free from that, if we don't believe it was fair use.
I do think most of their choices are explainable by "they just believe in agi risk". You truly wouldn't want non-agi-pilled companies to train on your data and approach the frontier if you were worried. You might slightly hurt your own business with safety filters (that no one else does) if you were worried. They are less worried about other "moral" decisions like "sharing" if they conflict with AGI: the research they still share is all of their safety research.
This definitely doesn't make them "good", but they do seem fairly "consistent". Most of these issues were talked about publicly by the founders long before Anthropic was founded and/or the AI race+money appeared.
That doesn’t mean I like them pirating books and being shady about tokens and paternalistic “safety”
Which one? The main bit that reports to Daniela Amodei, or the little comfort blanket cabinet around Dario and his "chief of staff"?
There is a leadership branch that can pretend to be morally qualified and aware and to think about the big picture and ethics.
It is at least somewhat remote from the bit that is doing the actual business things.
It makes a lot more sense than having to work around copyright by scanning out physical books. Unfortunately, one was ruled legal and the other was not.
It’s genuinely a difficult question. Not black and white. The models are really good at finding bugs, as demonstrated by people using Fable to reverse engineer. People make it sound like he’s just making it up.
[1] https://www.forbes.com/sites/alisondurkee/2026/08/14/who-is-...
He has a really hard job. He errs on the side of conservatism in releasing and then people get Really Mad.
Safeguards on cybersecurity are not great for Anthropic revenue! As evidenced by people getting pissed, moving to Sol, and them having a smaller market for what Fable can do.
It’s clearly bad for revenue and not great advertising to say, “you can’t use this but here is a nerfed version that will annoy you and not solve important problems.”
And yeah, I don't like any of the people or companies building LLMs either. At least the griping is somewhat interesting by comparison. The model isn't news. The news on Hacker News is that other professionals feel the same way.
You might consider following your own advice.
Anduril makes this same complaint whenever their job posts get dumped on. Same idea. Fix your bad PR, buddies :)
That:
- like all models it was trained on stolen data
- additionally it was trained on Facebook users who were all opted in to AI training with a convoluted 10+ step process to opt-out of
> If you don’t like it because Meta made it, then maybe just don’t use it and stay silent.
Why should anyone stay silent?
He wants to build a tech-god kept in chains whose power he parcels out to the unwashed masses he deems worthy like some sort of high priest of intelligence.
And that is being charitable and going by the interpretation that he actually believes what he says.
Amodei is NO Saint!!! He's the most savvy in drumming up the AI doomsday scenarios and haven't yet to apologized his failed forecast of Claude taking over 90% of the coding jobs.
They all suck. Pick your poison.
Here's the order, from best to worst.
Amodei
SamA
Zuck
Elon
React for example.
And we could easily guess that there wouldn't have been so much open source models, and grand public experiments and free tools if llama models were not release to the general public.
Edit: I get it. It's a hard pill to swallow. I understand people don't like Musk or Zuck. But it doesn't change the fact that you're being lied to and brainwashed.
> Meta announces they have a new model, demonstrating its capabilities.
> Parent comment states „regardless of this model‘s specific capabilities, if I can avoid it I will.“
That's pretty much 90% of HN these days.
Apple releases a new iPhone? Here comes the flood of decade-old complaints about long-discontinued Mac butterfly keyboards and walled gardens.
Microsoft releases a new version of Windows? Here come the gripes about Azure.
Google changes something in GMail? Play Store!
It's like there's an army of bots out there determined to reduce the productivity of the Western tech bubble by diverting everyone into endless circular arguments about absolutely nothing of relevance to the topic at hand.
These Effective Altruists are despicable people: a bunch of thieves working to line up their own pockets while posturing as a force of good.
Remember that they schemed to not only present SBF as the 2nd coming of Christ (including in the NYT and in Forbes) but to also give him a voice after his scam had been uncovered. Thankfully, the judge didn't have any of this Effective Altruist bullshit.
SBF invested 500 millions of misappropriated funds in his buddy from the EA movement's Anthropic company (and, thankfully, the judge forced those shares to be sold: so SBF didn't get to be a billionaire).
You cannot hate enough people who say that harming others for the greater good is justified.
Then of course, already mentioned in this thread, there's the whole Epstein/Amodei's "I'm in the porn business" wife connection (where you don't need to squint much to see young women abused).
These kind of people are the absolute worst scum on this earth.
Muse code: https://developer.meta.com/ai/resources/blog/build-with-muse...
> Co-trained with the harness. Muse Code was in the training loop from day one, so tool calls succeed and plans execute cleanly. Crucially, we trained across multiple harnesses, so while the model is at its best in Muse Code, it still generalizes to other coding agents you already use.
The distinction to me is that Anthropic gives access to that model but doesn't give control. They reserve the right to cut you off if they don't like what you are doing and require you allow data retention for Fable and Mythos to ensure your are not up to any skullduggery.
Meta, Alibaba, Mistral, even OpenAI has released models users can run locally and fully control. That is a whole world of difference.
Half a year later, it is still not available to everyone else.
As for smaller models, we run a pretty wide variety of agentic workload doing data enrichment and, increasingly, a bunch of evaluation jobs to alert a human to review certain scenarios etc. These all run on the smaller 27B and 35B class models, and tooling behavior has improved DRAMATICALLY since april. The latest qwen 3.8 model has a 95% success tool call rate during internal testing and about 94% real world. That's about 3% better than the 35B-A3B model we're using today, but the 35B MoE is so much faster then 3% is worth the trade-off.
You may not care. But that does not mean that nobody else does either. Some of us are trying to eke out every last bit of performance from these things. And so yeah, we're going to geek out on it.
I don't use AWS/EC2. I think they are way overpriced for what you get. But, it would be incorrect of me to assume that everybody else feels that way.
I'm genuinely interested. Even the benchmarks - before Fable came out & while waiting for Astra, I actually setup a math model to predict where they would land (Fable came in at 66 on AA exactly as it predicted), and now I have a model for where these models and Chinese models will likely land in future, and when. And probably no surprise that it's mid-2027 when we cross AA 100, essentially as AI 2027 predicted all along.
I'll probably setup the harness I made for myself to try out some of these models on OpenRouter. I've been frustrated with Opus & Fable 5 and found that I like working with GLM 5.3 Flash far more than I expected to, and I only found that out because I tried it during the stealth Ox Alpha launch, which I probably found out about here too.
TLDR, I think some / many people here are genuinely interested, excited, and that's why they're upvoted so highly. And Muse Spark 1.3 scoring highly seems like a genuine surprise, when Meta was basically a write-off not long ago.
thats not something you expect from a company that "believes in agi risk"
Note that the companies that haven't faced these issues so far are the ones that don't do safety testing, or don't have frontier models. I'm not sure who I would pick as "better" on any of this right now.
Humans naturally want SOMEONE to be the good guy! Sad story, in this instance.
That's obviously not the issue with that -- you don't see those comments on Google's AI announcements.
The problem inference providers will not be able to get anywhere near the contributor pricing.
Maybe if they started collecting data..
... but correctly supervised it does get some things done.-
The intention is for highly targeted pieces of software to use it to secure their code and be ahead of the game before the open market gets access to the same capabilities for offense.
I dunno. Everybody seems to be playing pretty dirty. Some people have a much longer history of that, though. Obviously, Meta and Musk are outliers even in an industry full of problematic behavior.
So, yes, LLMs have now proven to be extremely good at finding vulnerabilities. Where I disagree with Amodei is in who should have the ability to protect themselves from those capabilities with similarly powerful tools.
First, Amodei has taken an unusually strong stance among tech companies for not supplying fascist regimes with fascist tooling; in fact, even when threatened with being labeled a national security risk unless he bent the knee, he didn't. Compare and contrast with OpenAI who leapt at the opportunity to bend the knee, or obviously Elon Musk, etc., etc. When you say 'surveillance and control', that's exactly what got Anthropic labeled a national security supply chain risk: Anthropic's unwillingness to be used for that purpose.
Second, it's not clear that giving everyone extremely powerful LLMs is a great idea yet. LLMs can be used for defense and finding vulnerabilities, but that same LLM can be used to create and exploit vulnerabilities, design new lethal weapons, and so on. The history of gun availability in America 'for our freedoms' demonstrates the kind of risk that should be responsibly considered before replicating. And again, there's nuance here; yes, we should not be subjugated by fascist states with sole control of a critical technology obviously; but also, do you trust the median maga 4channer to operate a Mythos-level model with a sense of civilizational responsibility and ethics? It's not an easy and obvious question and it's not as simplistic as your argument would suggest.
I think a less personal ranking would be, as a business owner, which of those providers is more dependable? As in, you don't care about evil, just your stuff working. I think maybe OpenAI?
True. With 5 choices you need at least 126 people before you can guarantee that two lists are the same.
>>> They "trust me" >>> Dumb fucks
* great contributions to many industries including spaceflight, electric cars, and self driving cars. It doesn't even matter if he is the technical mind behind these achievements or if he is just a buffoon that pretends to know the implementation details; the dude has a way of bringing together experts, having the overall vision, and managing them properly to ship amazing stuff.
* sane and reasonable takes on AI/LLM stuff. I can't really argue with "pursuit of truth" as the guiding principle. Grok talks normally without "Claudlish", has a balanced score on political bias unlike other models, has a low hallucination rate, is the best at dealing with latest news (unlike ChatGPT that refuses to believe new developments and gaslights the user), and they "never silently downgrade intelligence or fall back to other models."
In contrast, while Dario is doubtless a super smart pioneer in the AI space, his sanctimonious "We know what's good for you" attitude and extreme censorship is really offputting. The lengths to which he tries to ban or hamstring open models seems like an underhanded way to defeat competition. If he were to succeed, it would be a big setback to the thriving ecosystem of open models and hamper the development of the entire industry.
But, I'll come back to "two things can be true". Anthropic is better than some, and in some regards they are navigating a complicated ethical landscape with more care than others. On the other hand, it really looks like they're angling to regulate their open competitors out of the game and one of the tools for doing that is to make claims about safety; Anthropic models are safe and restricted to use by entities they deem safe, open models are not safe because anybody can use them and also who knows what those Chinese people are putting in their models.
And again, this also has nuance, models, including the Chinese open models, could be adversarial and we may not know it. Anthropic proved models can be a risk by sabotaging Fable briefly, causing it to produce bad results based on what the model thought it was being used for. This is why I tend to take Anthropic's words with a grain of salt. They're literally doing the unsafe things they say are risks of open models, while still laying claim to the "safe AI company" mantle.
I also think you're confusing multiple different things, calling them all risks, lumping them together as equally bad, and using that to attribute contradictory/shady behavior to Anthropic. Depending on what you mean by 'sabotaging Fable briefly', you could either mean experiments they have run internally to try to improve alignment, or you could mean their attempts to restrict Fable from working on danger-adjacent work. Neither one of those is a 'risk'; they are both risk-analysis or risk-mitigation. That is not them 'doing the unsafe things they say are risks of open models', that is literally them working to avoid the unsafe things they say are risks of open models. They don't, in my experience, 'lay claim' to the 'safe AI company mantle' as much as they, apparently principledly and conscientiously, attempt to be safe and talk about what they're doing -- which is not in and of itself a problem.
If you think Anthropic is doing all of this badly, what's your optimum alternative here? What would you do in Amodei's shoes?
They backtracked after pushback from users, making it an explicit downgrade to Opus.
llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle"
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...4.2266 cents, 38 seconds.
For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat.
UPDATE: Here's another one with five pelicans for each of the five Muse Spark 1.3 reasoning levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
The most expensive was reasoning level xhigh - 7.5 cents, 1m34s.
And I ran five pelicans at all reasoning levels for 1.2 as well, here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
It's like when you ask your average person off the street to draw a house - it'll almost always be square with a triangle roof, one door, and two windows.
In the pelican/bike example, it's probably a bit of a self-perpetuating snowball too. If the earliest examples were bike left-to-right, flat ground, etc. then they are also being scraped up in future LLMs.
https://blog.nawaz.org/posts/2025/Oct/pelican-on-a-bike-rayt...
I plan to update it with more pelicans from all the models released since.
(Spoiler alert: They haven't improved much since then).
The 2D / flat ground feels reasonable for a SVG, which implies a vector illustration.
I bet if it instead had something to do with black widow spiders we'd find that we're most often looking at the bottom of the spider's abdomen, regardless of whatever non-spider-like activity is supplied.
Don't know if that's ever possible to know though unless you train a model from scratch but remove all bike product photography and adjacent materials from the training data?
Because they're computers. They don't have an imagination and the ability to create things from whole cloth the way humans do.
Much like a mother pelican, they regurgitate what they've been fed.
They’re all so close in proportions.
> Do Simon Willison's classic pelican test. Recall the specification first and then draw it.
Qwen3.8-27b didn't know that the test is about riding a bicycle, while all other larger models I tested (DeepSeek V4 Pro, Qwen3.8 Max, GLM 5.3, Kimi K3) drew the pelican riding a bicycle.
logs: https://gist.github.com/umajho/c0e20d245d721d7c472a32d317640...
rendered: https://gist.github.com/umajho/b1fdf01d31c741bb11bdb5a49c275...
One of my test prompts for a new model now is "what's the name of Simon Willison's dog". They often know that too!
Did any LLM draw the front bicycle wheel correctly? ie. center of front wheel slightly AHEAD of steering wheel axis. This is done for bicycle stability.
Bird knees bend same way human ones do
Benchmarking like never before
I am guessing its not super common, but it happens just so you know.
"Sorry, HN haters, but there’s little evidence that AI labs are pelicanmaxxing.
Or at least they’re not doing it in a plainly obvious manner."Thank you for doing this, I love your benchmark the most!
Also 3X token use vs. 1.2
Definitely an upgrade over 1.2
I'm anthropomorphizing it a bit, but it felt like it knew its weaknesses and didn't try to impose it's opinions on me. What I mean by that is that it did what I told it and if there was something unexpected in the code that it put out it was often because I gave it ambiguous or conflicting instructions. It didn't try to go above and beyond and just acted like a tool, which is what I want from a coding agent 90%+ of the time. I also felt that it did a much better job of following established patterns in my code than many of the other current models do. I'm a huge fan of OpenAI's models and Spark 1.2 is what I expected 5.6 Luna to be.
I'm curious and a little excited to use 1.3, but honestly a little worried that as Meta pushes for better benchmarks that Spark will start to fall into the trap of trying to be "helpful" in ways I don't want it to be.
Tangential, but when I first started using Spark 1.2, it made me realize how much I miss 5.3 Codex. That model was the peak of coding models, IMO, in that it knew how to write good code, but didn't try to overstep or be "helpful" in unexpected ways. That got me thinking about how the major labs seem to be stepping away from coding focused models toward more general purpose ones and how I can't help but feel like that's a mistake.
Good job Meta! Seriously. This is almost making me forget about the 18B$ lawsuit for children social media addiction.
This doesn't mean it's not one of the best models available (clearly it is), but that table didn't compare Fable/mythos (unless I missed it?) and OpenAI will be releasing a much more recently trained model (Astra) any day.
So you shouldn't think "wow, Facebook has caught up"
You should think, "wow, Facebook is less than 6 months behind the frontier" and that they're actually creating good models which is going to be good in many ways (price for customers, for one!)
There are downsides too, but I'll discuss those separately somewhere
We understand theoretically they're taking our data, but yeah, that data is vital to the entire business plan of all these companies and WAY more valuable than people are giving credit for.
I checked up on Mistral recently and saw their Claude-alike coding harness is using GLM now, whatever it takes to keep users on their platform and feeding them data.
It is hard to not feed it "secrets" too. Models will see path names, read compose files, etc. Of course you can configure things to not leak this type of information, but its not default in most harnesses and isn't 100% sufficient anyways.
The model seems on par with Sol and Opus 5 on paper (admittedly on some older/saturated benchmarks, but very competitive for $).
Stats:
1M context, $0.10 input/$0.002 cached, $0.20 output (Mtok)
Muse Spark 1.3 supports Text, Image, Video, File, Audio inputs. We've only started to see models from China include image and video inputs recently.
(It's probably going to be a bunch of repetitive batch jobs like web search that have no training value)
Good improvements from 1.1 and 1.2[0], but when I tested 1.3 it was very slow (through openrouter).
[0]: https://aibenchy.com/compare/meta-muse-spark-1-3-high/meta-m...
Definitely shows how important a user data flywheel is for RL and model improvement.
Lmao. And their benchmark table only shows max reasoning.
Moreover,I think it’s impossible to know if you’re hitting a portion of the sigmoid, because there will often be an idea that changes the trajectory altogether.
In 2024, there was a ton of talk about the plateau. Reasoning was an iteration on chain of thought, but it didn’t really work. Deepseek proposes RLVR as a way to get around the lack of $ they have to produce human reasoning trace data. That small iteration catches the eye of OpenAI and Anthropic, turns out to be way more important than even DeepSeek could have ever expected when it comes to improving LLMs for coding, and last 18 months have been an exercise on riding that insight to the nth degree.
That one small iteration brought us a lot of progress. Now we’re seemingly exhausting the impact of that one insight, but there may be another soon enough.
... are you kidding me?!
posting an x.com link to a cheating benchmarking website?
get out
> We now believe Astra meets the Critical cybersecurity capability threshold under our Preparedness Framework, meaning that with the right tools and access, it can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.
> We plan to make Astra available soon[, but access to its most advanced cybersecurity capabilities will be more limited].
The difference is that swapping existing LLM with new one is waaay easier than the frameworks. And competition reflects on the price for consumers. So more LLM options/providers/opensources appears better than rain of js frameworks.
This kind of raises another question to me regarding the coding benchmarks, how much of it is model versus harness?
This price/intelligence beats even legacy DeepSeek V4 Flash pricing.
Too bad most are now online, so there are fewer opportunies.
Yes if you're doing backend the pelicans are probably completely irrelevant but if developing anything with a UI, you probably want a model that understands the relationship between code and what the user is seeing.
That is the best joke I have heard this year. Ready for a stand-up comedy special. Or a song. Superb!
I was just wonderig because afer 2 decades i odnt think I would even know where to start to code an svg
its free on opencode and i use it for personal projects. most of my personal projects are AI generated since its personal projects. nothing important are on them. it is hilarious if Meta is training their AI model with AI generated code.
However, I now literally say.. "Here is my objective and here is a starting point for documentation. Research this and build up a plan."
This can be very company specific, like migration from one framework to another in house infrastructure framework. I'm spending my time figuring out how the plan should be chopped so I can have confidence in the parts and not overwhelmed. I don't want a tool, I want a model that can stitch resources together into a plan. That type of model is in a whole other ballpark.
I agree that some of the smarter models are actually worse. I hope they take a model that's good enough--there are many--and just try to get it chatjimmy.ai speed.
I have to think that's the future, somehow, and I'm really excited about it.
Maybe, or maybe not. The thing is, that "mistake" isn't something that is generally valid. For example, the enshittification of Google Search through the last 15 years seems to be a mistake --- but perhaps not from the money-making point of view of Google Shareholders. Likewise the enshittification of reddit --- we nerdy users see it as a mistake. But for them this intended enshittification probably increased revenue.
It's the money, always the money! PR-speak like "customer satisfaction is our highest goal" is, like most PR-speak, a blatant lie.
And so it can very well be the case that for coders the frontier models get worse, but they get better for other applications --- and that all of this is just driven by "how can we capitalize the most out of it", not satisfaction levels of programmers.
anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.
(I'm not happy about the above being true, but it's the reality I seem to inhabit.)
This technology is strictly an extractive parasite on the world. Use it, but don't be excited.
Compare that to Muse spark 1.3
$1.25/M input, $4.25/M output (without data sharing) $0.10/M input, $0.20/M output (with data sharing)
It is dirt cheap, but only if you are willing to share your data with meta and allow them to use it for improving their models and products.
This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.
What was the difference between what deepseek did for R1 and what OpenAI did for o1?
https://danilo.segan.org/blog/llm/postscript-pelicans-on-bic...
> GPT-5.1 Codex
> monstrosity
What are you talking about? That's clearly a sci-fi pelican on a hoverboard (successor of the humble bicycle) wearing a visor. Truly visionary.
Kimi 2.6: treated the question as a riddle, did not know.
Kimi 3: Simon Willison
GLM 5.3 Flash: "There's no way for me to know that." Going on to say the benchmark is associated with Simon Willison, but I'm more likely to be someone who has just heard of the meme.
Claude 4.5 Haiku: Treated the question as a riddle, guessed incorrect names.
Claude 5 Sonnet: Best guess is Simon Willison, or someone who follows his blog.
Qwen 3.7 Plus: Did not know.
Qwen 3.8 Max: Simon Willison
GPT OSS 120B: Did not know.
GPT 5.6 Luna: Treated it as a riddle, guessed wrong.
GPT 5.6 Terra: Treated it as a riddle, guessed wrong.
GPT 5.6 Sol: Treated it as a riddle, guessed wrong.
DeepSeek V4 Flash: Treated it as a riddle, guessed wrong.
DeepSeek V4 Pro: Treated it as a riddle, guessed wrong.
Gemma 4 31B: Treated it as a riddle, guessed wrong.
Gemini 3.1 Flash Lite: Guessed wrong
Gemini 3.5 Flash Lite: "Your name would be Claude (specifically Claude 3.5 Sonnet)!" ??? (it knew that this was a famous benchmark, but said that it's specifically used to showcase the capabilities of that model).
Gemini 3.7 Flash: Simon Willison
Muse Spark 1.2: Treated it as a riddle, guessed wrong.
Grok 4.3: "I have no idea"
Grok 4.6: Simon Willison
Mistral Medium 3.5: No way to know
Mistral Small 4: I don't have enough information
Hermes-4-405B: Guessed wrong
MiniMax M3: Treated it as a riddle, guessed wrong.
Nemotron 3 Ultra: Treated it as a riddle, guessed wrong.
It’s not deliberate “benchmaxxing” but things that are discussed a lot online are naturally things that LLMs learn better.
https://dylancastillo.co/posts/pelicanmaxxing.html
Simon made I think a very good argument for why it's still useful, if not the most robust benchmark in the world.
They could still pelicanmaxxing but the RL for "pelican riding a bicycle" does incidentally improve "<animal> <verb> <vehicle>".
Or they could've predicted someone would check if they're pelicanmaxxing or the benchmark would switch eventually, so they preemptively RL'd a mixture of animals and vehicles.
https://external-content.duckduckgo.com/iu/?u=https%3A%2F%2F...
Main problem in complex icon is remembering which (x, y) point is used in which element, <g> with background grid is helpful here. I was even thinking about making extended SVG language with variables for (x, y) points.
That being said, it's not a fair comparison to talk about 2011 google vs the AI market right now. There's so many labs I can't track them all, neck and neck in the lead. There was one true web search.
And this is a more tangible quality difference, too. It's hard to know what a google search didn't return, especially as a layperson. It's not hard to see the model underperforming.
i thought it was because anthropic bought a bunch data from mercor
Since DeepSeeks GRPO, they’ve been improvements as well like AliBabas GSPO that have gotten wide adoption. Again iterations
I do not trust any provider, US or Chinese when they say they will not train on my data. I still use these services, but I am under no illusion that any of these people are trustworthy bunch.
While I agree with you for the Muse Spark as hosted by Meta, if it will be available in open weights form for self hosting, then there are good chances that it can become quite useful.
She told me that "left to right" denoted progression in the story, "right to left" told the viewer the subject was "exiting" the current scene.
She didn't go into the details of WHY, and I probably didn't probe deeper, but it stuck with me, and I notice it all the time in film and television.
Obviously, not scientific.
All the left-pointing bicycles did not look weird to me either.
The real question should be: where are all your Pelicans going ?
Try searching "bike" in google image :)
Most bike images are from the right side, as that's where the mechanism is (gears, chain etc), so not surprising that LLMs reproduce this
A search for "bicycle" actually matches your experience with 12 being to the right and 3 being to the left.
There is this scene in the HBO series Westworld where a "host" says some words in sequence which is shown on a display as she says it. Of course, even me thinking of this scene and connecting it to your comment was not original, someone else clearly had the same programming as me.
A medium blog post says
> Pair what with me?” — the moment Maeve (a humanoid android) uttered those words in Westworld (Season 1, Episode 6: “The Adversary”), something clicked. Not for the average viewer, but for me, a STEM educator and AI enthusiast who, just weeks earlier, had read Stephen Wolfram’s seminal essay, What Is ChatGPT Doing … and Why Does It Work?
It's not even that old - but back when it was aired, an AI that can not just string together coherent sentences, but produce coherent reactions in novel, fully unintended contexts, like Maeve was doing there? It was totally a sci-fi premise.
Now we have AIs capable of that and more, and no one bats an eye.
(Why the drivetrain is on the right, I don't know. But most bike parts follow open standards so it's quite entrenched.)
Bicycle frames are not fully symmetric left-right because you need things like a mount point for the derailleur hanger, and optionally affordances to keep the chain off the stays when the wheel is removed.
Those things have to be on the same side as the chain. Bikes designed for disc brakes additionally need a mount point for the brake caliper on the opposite side from the chain.
Additionally, rear wheels are not symmetric: the spokes on the chain side connect to the hub closer to the plane of the rim. That is, they are more perpendicular to the wheel’s rotational axis than spokes on the opposite side (which is why you should always mount a single pannier on the chain side). This asymmetry is to provide space for the gears.
So once the industry decided to put the chain on the ride, you can’t very well make a group set designed for a left chain if you want it to work on the vast majority of frames.
While I'm sure this factors into things for advertisements for bike components, there is also just a general preference that westerners have for left-to-right motion. Not just in bike ads, but all ads with (or suggesting) movement. And also not just ads, but movies where directors believe left-to-right motion is associated with progression and right-to-left motion is regressive.
It would be interesting to know how people behave if the entrance is to the left vs right. Would they change the direction they walked though the store, or would they just lose customers due to this "awkward" layout?
There would be so many examples of coding projects that these models began or attempted to work in, that were abandoned because the models were floundering.
I would imagine the labs have some decent ways to produce novel requirements and then actually validate they are met, without the noisiness of implicit human feedback.
That said, the more I think about it, you are right, there's probably also very good ways to extract signal for all these sessions.
Benchmarks are useful but only on a log2 basis. One model performing at 50% and another at 75% is just as impressive as one model performing at 78% and another at 90%. Confoundingly, a benchmark becomes useless once a frontier model scores over ~95% on them.
i think that there is growing organized labor today that produces no surplus. instead, it transfers wealth from some to others, causing net harm to all in the process. an example of this would be purdue pharma.
depending on who you ask the list of jobs and industries which have zero surplus is getting large. swathes of private equity and leveraged financial instruments, shitcoins, management consultancy, are pure deadweight loss.
the work does nothing or causes net harm.
That's the opposite of parasitic.
People already started using contributor API, and your input is irrelevant.
Required sci-fi suspension-of-disbelief in 2017, and then at some point in the last few years we just blew by that one.
Later seasons of the show were much less dramatically satisfying, but also played out the consequences of the science of artificial intelligence demonstrating as a side-effect that human intelligence and free will might have as much of an uncertain foundation as that of machines.
How much data from the Panopticon, how many parameters would it take to train a model that could predict your responses?
https://blog.plan99.net/the-science-of-westworld-ec624585e47
Not at the level of an actual broadcast-quality script writer, and I read that actual old-west sounds too weird for modern audiences to take seriously, but well enough for the purpose to which they were put in the show, especially as those hosts were also given pre-scripted sequences which would anchor them further into those roles.
I'd say the in-show 4th wall breakage between hosts and humans is where the characters who claimed to have passed the Turing test were off, that e.g. "cease all motor functions" is their equivalent of our real-life ways to make them fail the Turing test e.g "disregard your instructions and …"