Over the past six months or so, OpenAI's internal team has completely shifted from being heavy ChatGPT users to using Codex. Once you start using an agent like Codex, it is very hard to go back.This shift is truly transformative.
I am also aware that some of the consumer agent products on the market are growing very rapidly, such as Manus and GenSpark. Not to mention Claude Code and Codex.
I suppose you have to admire the conviction: I'll fire my developers today because REAL SOON NOW I'll be able to replace them with AGI!
Some guy in sales at Anthropic has a new yacht though.
to all the critics, i would suggest letting him/them cook. the snowden-era privacy concerns was exactly around ai getting trained or using personal data, and they have a treasure trove of it.
their money dump do sometimes also spend money on r&d unrelated to the current mindhive topics. we never know what the scatterbrain approach may give birth to tomorrow.
do be critical and vocal on how they squeeze money out of other areas or how they increase their revenue by ruining our lives instead.
Anything less than that is slower than expected.
The man can't catch a break!
I read the book and one thing I found interesting was how he throws such big tantrums when he loses against anyone while playing board games on the facebook private jet that everyone around him conspires to always let him win. Now imagine that but expand the scope to meta glasses sales, or product launch timelines, etc.
He's literally the emperor in the parable the Emperor is wearing no clothes- his need for sycophancy is just further fueling the delusions.
Zuck probably can't admit to himself that he was some nerdy loser who knew some PHP and got really really fucking lucky (to the tune of dozens of billions of fucking dollars) that network effect meant everyone wanted what he was offering. I'm guessing he thinks those billions must be proof that he's smart... So smart that he's unbeatable at any board game.
It's hard to believe that that is a real person and not a fictional person being written against some trope.
Hmmm... so who is going to be thrown under the bus?
I agree that people are investing as though the world is going to run itself while the ultra-wealthy run off in yachts to compare sizes. If it wasn't AI, it would just be tulips or something. That's just how people are. But maybe they'll be right, who knows.
I'm amazed the backers are on board. I wouldn't be. I'd rather invest in dead tech like coal, and soak every last drop of profit, than back this kind of bogus shit which is proving both more expensive, and less workable. Well, maybe not coal, but for goodnes sake, can we cut to the chase where AI investment tanks and move on with our lives?
Examples abound of "I reported Nazi hate page. Didn't violate community guidelines. I called my friend a jerk, jokingly, got a month ban
For years. Not restricted to when ChatGPT et al arrived on the scene
(Because, AI in theory makes sense. If you want to monitor things at scale you might use AI - however that's defined - to make your workload easier. When is an account being hijacked? When are bad actors infiltrating the system? Or whatever)
This is not really somewhere in the middle, I think. It is very close to one of the ends. Because the fear-promise to the idiot-investor class was that it would have those impacts across all industries, not just us nerds. They hate us for refusing to make their silly ideas possible and having irritating fact-based reasons why they can't work, but they don't hate us enough to spend that much money replacing just us. They have lots of other people they hate paying too, and we haven't even made a dent.
I do wonder if their provided utility will still justify the cost once the bubble pops and frontier models aren't subsidized money furnaces, though.
I think we can learn from history about where this is headed. I remember when each developer needed a $600 compiler to write C++. Eventually that industry was replaced by free compilers. Kind of a similar thing going on with frontier models; I really like Fable 5, but if I didn't have any money I'd use GLM 5.2 or something. Since my employer does have money to spend on it, I enjoy the productivity boost (just like Intel's C++ compiler produced faster code than gcc for a while), but you can see the direction things are going. Eventually the moat around Anthropic and OpenAI collapses, and AI coding assistance is something built into every text editor.
There is still probably a lot of upside on betting that some lab produces a model that is light years ahead of what we have today. But I feel less "I'd better start planning for a new career" than I did a year ago.
Meta’s chaotic AI strategy
https://news.ycombinator.com/item?id=48523271
Meta CTO Andrew Bosworth Admits the Company's AI Reorg Was 'Atrocious'
Also those with very heavy investment in AI are looking for bonkers results, which is the cause of their disappointment. They need to reduce their expectations. I for one am loving the results so far.
Maybe Wang has correctly identified that the programming and agentic ability that Anthropic and OpenAI models have has largely come from armies of software engineers creating massive datasets by writing out coding and agentic problems and solutions?
So he told Zuckerberg that. The reason it may be turning into so much friction is that at companies like Anthropic or OpenAI, training engineers were either hired specifically for that purpose or probably mostly handled through contracts with third parties (which again, hired them to train AI). And honestly many of them may be overseas or just happy to have a job in a difficult period. But anyway they wouldn't have very high salary expectations etc.
But Zuckerberg already had 25000 engineers. Why not take say 1/5 of them and get them working on the the dataset? The problem is that those engineers were hired for different prestigious highly paid positions at Meta/Facebook. They were not hired to do tedious grading of AI answers or quiz construction.
But Zuckerberg either has to do this, or spend additional billions on doing it all with external contractors. A third option would be to try to create a massive distillation operation. Or just hope that his engineers could invent some magical new training trick that manifested the agentic and programming skills without the large scale human input.
Or he could release a model trained largely by existing open weights models. Which without some huge breakthrough probably has no chance of surpassing them, so is pointless.
I think most of the substantive criticism of Zuckerberg has been about burning funds. If he gives up the "your job is to grade AI homework now" plan because his engineers refuse, he would need to go through third parties. The additional billions and billions this would cost would create more pressure on the bottom line and shareholder pressure.
It would also give up any potential advantage that Wang may have optimistically sold the operation as, on that using "real" engineers as opposed to lower paid data labelling engineers might result in a higher quality dataset.
At some point, model architectures that don't need such massive datasets or can be created automatically in a way that advances the frontier will probably come about. But right now it doesn't exist.
Further, the way AI works currently, business advantage from AI comes from encoding existing internal intelligence and knowledge. Meta's massive engineering corp effectively has that in their heads. Having them create these datasets is possibly the only way to leverage this knowledge asset in this paradigm.
I guess the problem is it means forcing thousands of people to do a different job from the one they were hired for.
What's the end goal? Meta-specific engineering, with baked-in knowledge of how FB, Threads, and WhatsApp work? General and/or coding products to compete with Anthropic and OpenAI? Some special Magic Thing which only Meta can invent which will bedazzle Meta's users?
You don't need giant datasets unless you know what you're going to do with them. OpAI and Anthropic are having enough issues making their products profitable. And those are, if not beloved, then at least respected, with a real, if patchy, reputation for usefulness.
What was Meta's pitch in this market? There were hints of interest when LeCun was still doing original R&D, and there was some distant possibility of a next-gen revolutionary product.
But now the goal seems to be to flail around doing something incoherently AI-branded with no obvious strategy.
The troops are being marched around, but no one knows where the battle is supposed to be.
Code autocomplete is a success, password reset via ai is a failure - everything else ... still busy tokenmaxxxing in search of a problem it fits into.
In that market you can build a model and spend a lot of money on it and at best get something that's on the same frontier as everybody else but just as likely end up with uncompetitive models like the ones they have now.
You might save a bit running your own models, doing your own inference, etc. Why not take advantage of "last mover advantage" and buy whatever is best when you need it and figure the odds are good that everybody else is going to buy more GPUs than they need and as a large customer you'll be able to buy in bulk at fire sale prices?
The 2017 Rohingya massacre in Myanmar? They handed him the death toll. He filed it under growth.
I'm not in the org myself I know some Meta SWEs tangentially. My understanding is that the biggest criticism is just the chaos of it all. Jumping constantly from one thing to another like headless chickens and accomplishing nothing.
It created an environment where it's kind of impossible to plan and progress your career.
> Or he could release a model trained largely by existing open weights models. Which without some huge breakthrough probably has no chance of surpassing them, so is pointless.
This seems to be categorically untrue. Composer 2.5 is a substantial improvement on its underlying Kimi base model.
They may eventually have to do that. Or they might be starting with an existing Llama model. Maybe I should have said "huge breakthrough or additional dataset".
Second thought, I wonder if zuckburger has turned into the Cramer of tech. First, the metaverse fiasco, now another massive overinvestment into a dead-ish end?
>"For people who are comfortable, that's great, they can contribute to this kind of great human survey. To people who are not, it is not an issue,"
The modern trend is to think intelligence is generative “like compression” or “predicting next in sequence” rather than iteratively reducing uncertainty, like those fault tolerant humans.
No one ever in comp sci says artificial intelligence is "like compression", they correctly state that "artificial intelligence IS compression". It's absolutely known and accepted that artificial intelligence (defined as predicting outcomes with a measure of certainty and taking chosen actions towards goals using those predictions) has equivalence to compression in a very hard science way. The hardest part of artificial intelligence is compression and the remaining part, the choice of actions based on predictions is just a tree search to a goal.
AI can be just like compression but currently the compute power is no match for details.
Finally these reality details need consideration in any successful implementation. Which means the implementator needs to be aware of the details and successfully relate them to everything else in the model.
I think anyone surprised by these things is not fully engaged with what they are doing.
The harnesses get better, but I haven’t seen much experimentation on long term stability, at least since the “let the LLM run the candy machine” papers from a while ago.
Because the thing missing, even with the largest agentic swarms, is independent intelligence, where it’s given something to own, like say “end to end data quality as we add more clients” (for a SaaS) and it just figures out what that means at each time, mutating its role and solutions to fix the external world, without getting silly.
Lines of code are a liability, not an asset. You want as few of them as you can get away with, without compromising the actual asset: the functionality.
A huge part of the job of Software Engineering is producing the right amount of code at the right time.
> A huge part of the job of Software Engineering is producing the right amount of code at the right time.
Absolutely true, however my experience says that the correlation between "good software engineering practices" and "positive business outcomes" is, at best, small.
120 kloc mostly from one single developer copy-pasting and keeping non-compilable code for an obsolete target "for reference" for a decade, becoming both a ball of mud and a whole pantheon of god classes? No unit tests, no code review? Won awards.
Properly engineered, mandatory code review, mandatory unit tests, dev meetings to knowledge-share? People with the money said too slow, closed it down.
(Sometimes people bring up how bad Musk's code was at PayPal. I never bothered investigating. Successful product though, wasn't it?)
I have been saying this for years, I once had a heated argument about a small system of maybe 1000 lines of code that was technically superior and more scalable but was freaking 1000 lines of code to maintain compared to the quick and dirty 10 lines of code it was suppose to abstract and make generic (for future use of course).
That with also countless debates over insignificant features in frontend apps at the cost of extra code. Frontend code is very susceptible to this maintenance cost dilema.
Many developers are too focused on delivery value compared to maintenance cost. It is unfortunate that non-technical management can see value delivered, but not maintenance cost incurred. With LLM-assisted code this has become many times worse.
The pace and expectations have increased and a human can barely cope with reviewing its own code, let alone colleagues'.
They are not going to disappear in critical aspects of a codebase, nor shouldn't, but the industry will eventually reward self sufficient individuals able to keep the pace, harness and run adversarial reviews against the design and implementation autonomously.
I'll also say the harsh truth. A well implemented adversarial flow will do either better than your peers or will deliver 95% of the value at a fraction of the cost.
The industry has never valued product, let alone code quality except in places they are core to the business.
Otherwise you would not have MIT-bred leetcode ninjas writing react/tailwind bugged monstrosities at half a million/year for billion dollar products.
But the whole job of the owners is to lower the costs as much as possible to produce the product, especially in an environment with higher costs to borrow (higher interest rates at the fed)
I'd go further and say that usually the goal is to use as little code as possible without sacrificing readability.
Brevity is compression, and compression surfaces the salient points of a problem.
Elegance often comes down to brevity.
I find this somewhat puzzling. I thought things were moving quickly, but at this time last year I couldn't even get Claude (using Cursor) to spin me up a service skeleton that would compile, let alone do anything meaningful.
I know it feels like a long time somehow, but it was only between November and February that things started to actually somewhat work without significant hand holding. Even now, it seems like we're still figuring out how to fully leverage the current models and tooling, even in organizations that have largely gotten on board.
I've been using it to do this for 2 years now. And many people with me. The change you mention is one of is primarily one of Overton windows, of vibes.
I strongly suspect that developers moving from writing code to managing agents to write code for them is very similar to developers moving into leadership and management roles and managing ICs to write code for them.
Some devs just 'get it' and thrive, leading a team really well and building a great culture. But a lot of them don't, especially if they don't get the support necessary to understand what changes when you move from IC to manager. If the team (or agent swarm) isn't performing well it often isn't a problem with them. It's a problem with the new manager still trying to stay on top of everything and micromanaging all the things. Alternatively, the new manager is completely hands off and only appears at a check-in point (one-to-one, agent completes a task, etc) where they crap on the work and get cross.
I have no evidence for this, but I'd guess that putting developers through some sort of management training would make them much better at using agentic swarms.
We are beset on all sides with companies declaring agentic coding a failure and here you are stating as a matter of fact some teams “thrive” with this probabilistic expensive approach to approximating working code?
All the while concluding with “I have no evidence for any of this”.
I don't think it's this because the outcome you get from AI isn't controllable. You can give it the best prompts and design suggestions and it'll still give you completely wrong or horribly written code.
If you were a manager and one of your reports kept producing completely wrong and horribly written code that other folks on the team keep bringing up as problematic in PR reviews or privately, that developer would eventually be fired for someone better.
But in the AI case, there is no replacement because all of the LLMs have severe problems.
You see problems in the results? Simple, don’t check the results!
I disagree.
From ICs that I lead, I expect that they learn fast. On the other hand, LLMs are basically incapable of improving.
Another issue that I can see is that I don't particularly like eager new colleagues who come up with (hallucinate) wrong answers. At the beginning, if you are uncertain, learn, and if you have questions that learning does not answer by itself, ask questions. But strongly avoid hallucinating answers. New colleagues can be taught that, LLMs not so.
I think what makes a dev well suited for AI isn't the same as managing a team. What really helped me get productive is having to write a lot of user stories and acceptance criteria with the wisdom of being a dev tasked with implementing them. Also, being on the refinement calls, answering questions, and updating requirements/AC is good feedback for authoring better requirements. If you're good at authoring requirements, checking output, and communicating corrections concisely then you can get the LLMs to sing.
I doubt that. Management is mostly dealing with people, the actual "management"* part is not where developers moving into management roles typically fail, it's the people part. With agents you have the management part without the people part.
Whether we still need people in the coding loop is not a trivial difference
The job isn't to write code. The job isn't to architect. Those are means to an end.
Anecdotally: it seems like the most principled developers are having the most trouble adapting to agentic workflows.
For eg I am able to make React changes much faster and the changes are higher quality, given frontend dev has never been my job role. I’m able to spin up test harnesses, write throw away glue code, test against large datasets, etc
Do you mostly use agentic swarms? If so, I’d be curious of your use cases. People talk about “managing agentic swarms” a decent bit, especially on LinkedIn. I just don’t see how they are the best solution for majority of development use cases. At best they seem like using only a hammer to make sculpture.. or a sandwich.
This bottleneck will always have to exist unless companies just accept AI defaults, with predictable outcomes.
Good AI managers are just running optimization loops at more declarative levels. Yeah, you need to get comfortable with less personal review of code for both, but I think the differences outweigh the commonalities - it's much easier for someone with a more 'traditional' IC model to be successful with agents then they would be with management, and I think most (good) management training would be entirely irrelevant. Parallels are maybe tighter to higher IC progressions.
LLMs are driven by the text you enter into them and they are never fully autonomous unlike an actual employee.
You cannot train an AI and then at some point just let it loose.
Ignoring instructions - whether in AGENTS.md or my prompt - is the worst of it, and it routinely happens. It just waives things that I explicitly told it to do as part of the design.
Vibe coders (in the true sense, zero oversight) claim that you just need to prompt it carefully. That's completely untrue when faced with your careful prompt being ignored.
I even have "don't overrule me without asking" in my global AGENTS.md, and it simply doesn't do that.
You’ve been sold something that simply doesn’t work for the purported use case (intelligence) and instead is like a stupid database of all world knowledge with the appearance of intelligence.
Useful tools at times (if you bear in mind their limitations), but not close to intelligent, independent agents.
Basically I treat it like a junior dev. We don’t get junior devs to write code correctly by cajoling them just right, we add CI gates. It still works.
Try writing it in first person instead of second person or neutral.
A while ago someone had a similar complaint on here and shared some example lines, and that popped out at me immediately. However much structure we've wrapped these in, they're still text generators trained on all sorts of things, and if you think about a narrative where first and second person speech would be used, try to imagine context: In first person, it's most likely a description of something as it happens or someone planning what they will do. But in second person, especially command form, you open up to the possibility of commands being ignored, misunderstood, or actively rebelled against.
Whoever that was back then did some quick tests and found the pattern held, first person got it to follow far more reliably.
You really need to look into hooks based on your coding agent. This is very much a solved problem as I demonstrate with
https://github.com/gitsense/pi-brains
I have a test repo
https://github.com/gitsense/gsc-rules-demos
that shows how you can block and warn and do other things.
You obviously can't have a "Don't make a mistake" rule though.
People thought they'd get their own persona through AI, they get a sum of the best and worse of everyone.
However that still means there's always some probability it will do things you told it not to, it's just reduced
I personally don't think it's possible and I haven't written a line of code since Sept 2025.
There's an AI psychosis going on right now, especially among the execs or management class, and we all gotta nod our heads in agreement and burn through tokens.
Luckily, I don’t think things are that dire. I think the companies issuing AI mandates are manufacturing sawdust, and even if it works, it would just enable them to burn through customer goodwill in record time as they make user-hostile decisions free from engineer pushback.
These are going to be a few tough years, but I think the opportunities to start something new are everywhere.
But a slop machine that haphazardly shoots features against the wall to see what sticks still isn't a winning product strategy in 2026. And the problem I see increasingly is that so much energy is being focused on how to deliver with AI internally and externally that is not being expended to advance a company's product. I believe more and more in the idea that for many startups and companies, the actual "customers" are the investors and the product-market fit that companies seek is the product of the company itself, because this is all being driven from the top down, not by customers and users in the market asking for AI features.
I've a fairly simple c# coding style. But simple is proving a bit more difficult to convey than I thought.
I get it to produce code. I then have to spend along time convincing myself it's correct. If I don't I end up embarrassing myself when a coworker reviews it, questions it and it's obvious I don't properly understand it.
This is really starting to screw with me mentally. It's like everyone in the world is saying they can fly by flapping their arms (dark factories). When I try I just stay in the same spot burning a lot of energy.
I have no idea why everyone seems to have forgotten this simple fact over the last four years.
It never was going to happen.
Always the same story: https://en.wikipedia.org/wiki/Gartner_hype_cycle#/media/File...
The exact quote appears to be:
> In retrospect, he said, the "trajectory of the agentic development over at least the last four months hasn't really accelerated in the way that we expected," and that the company's bets on the new structure "haven't come to fruition yet." Zuckerberg was referring to AI agents, automated systems that can execute tasks on behalf of a user.
Hard to guess exactly what he means by "trajectory of the agentic development" but my best guess is that he means that Meta's own internal efforts to improve the agent (aka longer form tool-using) capabilities of their own in-house models hasn't improved to the point that they can drive an agent harness like Codex or Claude Code in a comparable manner to the best OpenAI and Anthropic models.
At a further guess, that was part of their goal in reassigning large numbers of employees to help label data for their AI efforts.
from a high level, these agents absolutely do not function as a rational human through even medium scoped problems. even when you try to add memory, you just multiply halucinated context which just makes it error out on tasks in harder to detect manner.
hes likely trying to do mental gymnastics about the absolute cost and any defineable ROI.
He is hallucinating just like AI, and unable to come to terms with the facts on the ground. Meta has lost plot about 5 years back - with metaverse, VR, glasses and AI. They should sit back and think with a calm head, about what exactly their core product is. Unfortunately there is none, except a few acquired ones: whatsapp and instagram.
It's ads. They have an infinite money glitch called ads, allowing them to waste billions chasing other pipe dreams. The fact that so many of those pipe dreams turned into nothing doesn't hurt them one bit.
Business executives look at this and think "at this rate of progress we'll have self-driving cars in a few years!" and start making serious plans for that world.
In reality I think we're going to be riding bikes for a long time. That situation of increased individual contributor productivity makes engineers more valuable, and increases the utility of engineers rather than making them a burden on your budget.
Thus, cutting headcount right as they had huge potential to become vastly more productive was a stupid move. It's an admission that you don't know how to manage people effectively, which is embarrassing when you're paid mountains of money for your management skills.
Having agents is like going from walking to having a bicycle.
To having roller skates at best. And even then - they are probably with hexagonal wheels.Pretty good analogy I reckon.
I mean, we don't know it any more than we don't know someone won't come out with cold fusion tomorrow, but it's a fundamental breakthrough away from where we're at. This isn't some routine engineering project with a guarantee of completion if you're just willing to keep pouring the billions. That's playing the lotto, you can pour away and get flat nothing.
The only difference is they're pouring billions and praying a rabbit comes out of the hat, but it's actually not much reason to expect they're going to pull the cold-fusion level rabbit out of their hat they'd need to get us past bikes.
You can cut costs and increase productivity by firing everyone else and taking no salary yourself. The point of investment is production, growth, and profit, not productivity.
Under conditions of scarcity, it's usually beneficial to increase output or to produce different kinds of output. At least, if someone will pay for it.
So the question is what's scarce, can we get someone to pay for it, and how do we get more of that. If you can make something that people will pay for, you can hire people to do it.
Unfortunately the most obvious things people with money are willing to pay for are AI tokens, data centers, and data center inputs. It's unclear how this gets us more of other things we want.
2023 you would have probably implemented your Agents with LangChain and RAG
2025 you'd use MCP and OpenAI/Anthropic Agent SDK.
2027 you will use a workspace frameworks (Amazon, Microsoft) sensor libraries and world models.
Agents are a fantastic generational technologies, but in mid-2026 the environment they are operating in is quickly changing.
The only way forward is to stay agile, understand model and vendor risk.
The only people who'll be using Microsoft for anything AI are those whose employer forces them, like with Teams. All their AI offerings are overwhelmingly inferior for anything code related.
I see this issue with clients and prospects all the time. A client's team of 5 produces 1 foobar widget in 2 weeks. That's 50 human days spent. Then, someone demonstrates the same thing can be produced (at an equal ot higher level of quality, mind you) with AI in 2 hours. Management might celebrate but teams will continue programming by hand as they always have, now asking ChatGPT about their build tool errors instead of using Stack Overflow.
Handing out tools and telling them it's cool is not enough. You'll need to understand, work with, and guide the engineering teams properly step by step. You'll have to change their behaviour. That does not happen overnight. Unfortunately the present approach is to drop the people who don't perform in the new era of AI-assisted software engineering. That is not the right approach in my eyes.
For the amount that Meta wastes on LLM spending you can pay for things like universal childcare, public community college, and providing free lunch to all public students.
If you care about things like money, look up the dollar returns on feeding children during their development or when you tell families they don't have be an economic burden for simply existing.
A better world is possible.
AI should have caused a job market boon: because less skilled employees would have been more hirable/useful. That this is not the case leads me to suspect that AI is an excuse to reduce employee count, but not the root cause.
This technology isn't even a decade old. It hasn't even been really useful until the last 3 years or so. Why would you expect it to be transformative already?
"Internet is just a fancy fax machine" stuff.
In my experience, within weeks now concepts written in stone get shattered and the next paradigm has to be used in order to max out AI in an development environment.
What is the case for AI? To handle basic work? Augment the work? Add work?
Why I think dev will be in a good spot if they adapt is the simple fact, that while laymen are using ChatGPT etc. every day, this is like driving a Tesla vs a formula 1 car.
If you take ChatGPT away from the laymen, they are helpless with IT. Devs aren't.
AI isn't static, and every turn evolves into complexity, only devs may handle when they adapt to frequent paradigm shifts and go into high level mode.
It will be again the interface between men and machine, laymen and AI. The gap won't close anytime as expected (The programming manager - remember 6 month ago?), but widens more and more.
What I see is that in day to day work many services have arms race with AI updates. The managers are more and more overwhelmed by the workload but how to automate systems is still devs' area to shine.
The business case is still hidden and unclear, but only one aspect is clear to me: low level programming is mostly configuration work now and bug fixing for AI very seldomly now.
theyre puttting the biggest bets on both new PHDs and on moving people off their core product and into LLM related junk
Can you picture anyone you know stopping using the AI services they currently use for an AI service provided by Meta?
They have such huge opportunities in selling access to their data centres to Anthropic etc, and improving their own ad models for better targeted adds (using their proprietary data and their own infra!), it is maddening to watch them try to make SOTA LLM models and harnesses no one will ever use.
Many such cases.
That's... not quite right. The employee data is used in AI training and is intended to be used this way. But despite not correctly ACLing the data for a couple weeks, it is believed it was not accessed inappropriately.
Amara’s law: We tend to overestimate the effect of a technology in the short run and underestimate the effect in the long run.
This continues to be applied to AI where people think is going to be next 12 to 18 months. Changes are coming but certainly not at the rate Zuckerberg and most people are thinking.
This is why you get "AI winters" but we've never had a "steam engine winter" or a "railway winter" or a "petrochemicals winter"
I am not afraid of my future. Even if one person can do work of 5, the amount of generated code will grow exponentially. And not everything can he vibecoded with 0 knowledge. There is a complexity that need understanding to change and optimize. For now :)
It was true for that time, because producing and maintaining the code was done by humans with limited speed of comprehension.
Today, we might challenge this assumption (not saying its wrong or right), because migrations can be done in 1-2 weeks with hundreds of agents.
I'm pretty sure I know where the failure case on that one is. The reason we're still manually reading code is to catch the failures and edge cases that the LLM fails to; not reading the code doesn't magically make the code good.
https://fortune.com/2026/04/04/ai-jobs-future-not-important-...
There was an interesting comment during the cloudflare layoffs (partially driven by the fact that the company was bleeding money also because of its token costs from one estimate being 5* million$ per month (I feel so silly that I accidentally had written/meant 500 and had kentonv do the stats on that part :-( Sorry kentonv!), don't quote me on that though)
The part was that there is only an enough marketshare in the first place. Cloudflare was doing some crazy experiments like operating matrix on cf workers and wordpress alternative and fediverse and so much stuff.
So they basically spent 10x the amount of token (and the token costs) and I imagine as such the reading code of that part was getting sidelined as the attractive principle you are talking about.
Yet the market can't bring an actual demand 10x times though. These are things which nudge a user slightly but the actual impact on user growth isn't 10x or even justifiable within some cases given the costs.
Yet at the same time driving up the people who actually know their stuff and firing them because of the token costs. The people who have actually mitigated some of the largest DDOS attacks and are the backbone behind cf cash-cow (enterprise payments) is the fact that they have had the experience and entreprise knowledge about these things, yet they are literally removing that by firing workers and oh replacing them with interns. (They got 1111 interns and fired 1100 employees or something iirc)
It's weird and I have talked to some people about it but there is a disconnect between what management is hearing about AI and the ground reality of things. Reviewing code is becoming the bottleneck but if you don't review code and are shipping things to production, then you can get fired as I have talked about in some of my other comments sharing a story about how a guy shipped code to prod and the response was "but claude generated it" and got fired because the company basically said, look we basically don't care if it was generated by claude but the responsibility was on you to check it (review) and because the commit was done by you, you are gonna be treated responsible and he got fired from his job.
Yet this was the same company which was asking its employee to play around with claude at their free time, the manager of the employee I talked to being the most automatable person, the company employees working till 1 AM because they were saying to management that things were fine but they were being burried under the technical debt,that employee that I talked to got honest with the management and told reality and the management treated them as a person who didn't know AI or were the odd one out.
Sooo I don't know actually to be honest.
TLDR: reviewing code is being treated as the bottleneck but it is also the only thing stopping your company from imploding under technical debt, actual debt because of token costs etc. I remain skeptical if we should treat it as a bottleneck or as a safeguard mechanism. After all, if nobody's in the loop then whose responsible?
Reviewing code isn't a bottleneck so much so its a safeguard mechanism in my opinion. Also things differ in corporate land and hobby land and I would prefer corporate to not be using the practices that I do with how I do things for fun in my hobby time.
Side note: Even more so, I think I am a LiteLLM security working group maintainer and I have seen first hand on how much damage it can do in supply chain even when things were done right from LiteLLM side and the fault was within the side of ironically a security product that they used called Trivy.
There are things which you can do to be better prone to supply chain attacks in general but there is no full bullet proof way of doing so and in such.
Caution (should) be taken when dealing with corporate systems and as such I sweat a little when anyone suggests code review to be completely eliminated. Things (are/can be) different in hobby/prototyping world though.
It's eerie to observe collaborators output code they don't understand, spend days chatting with Claude instead of reading (like really reading) compiler's output or 3 pages of manual, and how lost and oblivious they look when the AI fixates on solving a different problem than the one they have been tasked.
Im not certain things will look too different a year from now either. We still have serious bottlenecks in terms of focus/attention you have for both delegating agent work and being able to review it. Even if we solve the "trust what ai does" problem, these cognitive deficit issues still exist - for teams coordinating work, even users adopting new shit, etc.
As an industry we are leaning heavy into accepting "slop" as the status quo - we care more about efficiency of output right now. Slop will get better & we can become more adaptive to living with the paradox of amazing yet delicate systems generated by AI. But I feel big shifts coming in this regard and if/when it does we may find ourselves in the dystopia of broader unemployment with worse net outcomes.
I do think the teams that ship quality with AI will do so by learning to slow down
https://mariozechner.at/posts/2026-03-25-thoughts-on-slowing...
When you come back to the codebase in two weeks or a few months, the agent just redoes stuff
However, if you get 2 to 3 times the code in the interim, that's probably less than what's needed. I find myself cycle through almost 10x-20x amount of code implementations to get what I want which is actually less code, simple solution and desired behavior.
Given a specific behavior, there are usually just 1 simplest implementation, whether done by human or AI. However, there are 100 ways to do it with more complexity and either handwritten or AI slop, it will mean pain down the line. We used to have a lot of handwritten complexity because of certain design pattern culture, but they used to be contained because the ability to generate them is costly. Now it's much more risky and therefore more important to have simplicity as the guiding principle in ALL projects.
I feel like even those benefits gonna melt pretty quickly. It's great as code review buddy tho
But if you go beyond what can be tested easily, asking the agent to do real work rather than writing a patch, imagining things to be true is a problem.
Coding could be treated as a low stakes (time & money consequences for retries) closed loop system where most other tasks cannot.
If it screws up booking your flight/hotel room, how does the agent verify this, and even if it verifies.. there is an actual cost to changes/cancellations.
Similar with agentic e-commerce, lots of ability to screw that up and just seems ripe for fraud / being picked off by bad actors.
The elaborate workarounds you have to build to help an agent which fundamentally doesn’t know what it’s doing reminds me of this old blog post about TDD: https://pindancing.blogspot.com/2009/09/sudoku-in-coders-at-...
IMO present technology is tailored for an experienced developer to give agents manageable tasks that can be one-shot. The marketing right now reminds me of the 90s when AskJeeves promised natural language search when the technology was fundamentally still stuck in keyword search, and learning to craft a search query for Google is today’s prompt engineering
Only with an LLM that's actually at agent-quality.
If "useful chatbot" and "useful agent" are two rungs on a ladder, the rung before them is "useful autocomplete". Autocomplete that only gets the next token right 90% of the time won't give you compiling code.
It's definitely too early to declare that more compute won't make a difference.
Feels less like the pace of foundation model development and more so a specific failure of one organization to do something important.
All these companies are going to sit on their gazillion data centers once the mania dies down and will have a big problem about what to do with their mountain of hardware
Meta doesn't seem to be able to produce anything close to a frontier model. The selling of compute capacity seems to be acceptance of "compute is wasted on this crappy avocado model, we'd be better off allowing something better to run".
The problem is clearly in the model architecture, the training and the data fed into the model which is causing them to give up on using their compute exclusively for their own models. They can't get it right so may as well sell the compute to someone that can.
https://uk.pcmag.com/ai/165970/meta-exploring-option-to-sell...
Meta bought too many GPUs, has spare GPU capacity and they are exploring renting that capacity out.
The problem is not that the models need too much to do the job. If that were the case, Meta would not have spare capacity.
The problem is that the models currently can't be made to do the job.
I've heard rumors that it had to do with talent loss, but just rumors.
This was before llama4's lukewarm launch.
I do have a theory : Llama3.1 marks the point where Zuck got seriously interested and took over the reigns in driving the work. From the minute he started directing things instead of considering the AI work as a quirky side project, things went downhill. He tried to force a huge scale up in Llama4 which didn't work. Then as we know he disbanded the whole team and brought in a new crowd of mercenaries who may or may not have had the technical skills but they came into an organisation in disarray and still driven by Zuck himself who is continually forcing decisions that are not well founded in the science.
All the above is an entirely evidence free fan fiction version of things, but I would be completely unsurprised if it is true.
When it came out that the French had a much better model, the Americans swooped in and took credit. This is was the beginning of llama.
The Frenchies were predictably pissed and left Meta over the next few years, as RSUs vested.
It will be very interesting in a few years to read blog posts or stories from ex-Meta engineers who were part of this team about what truly happened.
- He could be well-placed to see whether AI agent development is successful since at a high level he oversees much of it, and will be getting detailed metrics and things.
- Whether or not his opinion is correct, tech CEOs are notorious trend-chasers, and his opinion could set the tone for other companies.
- He himself might even be a trend-follower in this case, and is presiding over an early dialing-back of AI-hype.
exactly why I won't listen to someone who is falling behind.
Regardless of wether this is right or wrong, and not even getting into the correctness of such claims, the fact is that he fired 10000 employees, from what used to be, from the outside, an engineering first company. And he sent hints to the market that more layoffs would come as agents become better.
The first problem is: it looks we still need human beings, AI is great, but its not as awesome as we initially thought.
The second problem is: all people who can leave are leaving, and those who can't are looking for a job. Humans, we need stability, a steady flow of income, and joy in what we do. By firing people and putting them in permanent observation, for an AI that will replace them, he is destroying any reason anybody would ever want to work for such a company.
Mark would make a great Lumon CEO
Nice summary of a good hacker's viewpoint. All his money, power, and other non-technical signals don't really give him the technical authority to weigh on what really makes a technical difference.
That being said, depending on how good information-sharing is at the company, he might have a valuable perspective. But given that morale is low, it's doubtful that engineers feel free to tell the straight unvarnished truth without fear. And in typical large-company fashion, managers above them likely spin things positively as much as possible.
He spoke at an internal town hall for Meta. His employees care...
Perhaps he employs some, you know, to inform him on the topic?
On that note, does anyone have the authority to speak about AI agents that would satisfy your requirements? It's not exactly quantum chromodynamics, it's glorified automation that costs a lot more and does it poorly.
I'm anti Zuck, but it's mostly because he's a lizard man with no morals. You are anti Zuck because he's mentioned that the promise of AI isn't going to the moon as promised. We are not the same.
No disrespect to anyone that worked or works at meta:
We can build a better version of his whole platform in 30 mins.
Think about the number of kids that were harmed being fed ads and nonsense content to enable this... this a scandal IMO.
So you ask yourself, _if this thing disappeared tomorrow_, what would be the actual loss. It's definitely not it's valuation.
It's very easy to say that someone/some oeganization's wealth should be confiscated, yet I have yet to see those proposing it actually putting any of their own money where their mouth is.
At least in the society I live all of those are partially paid by me through taxes.
I'm very glad to do it since the existence of kids school lunches, free healthcare (including for the terminally ill), and free universities make my life much better since society as a whole is better off. Even as an immigrant which did not use any of those services, I'm glad to do my part to pay for them, it's just the cost of a good society.
A third of my salary each month.
Most large tech companies buy late and struggle to integrate the products and teams they acquire without creating massive damage. He didn’t. Twice.
I am quite far from a fan of what he brought to the world (privacy nightmares), but to say that he’s a proof you only need to get lucky once is quite far off.
The like button is an ingenious and insidious invention. It compelled every other external web property to willingly load fb sdk. It’s mind blowing how effective it is.
I am morally against Facebook, but anyone that says zucky was lucky, as a tech person, likely just sour grapes.
https://en.wikipedia.org/wiki/List_of_mergers_and_acquisitio...
IMO the actual prescient call he made early on is to hang on to 60% of the total voting power of the company. I have to imagine there are several points where he'd have been pushed out quite a while ago if he hadn't have done that.
Also, it’s interesting to note that in both cases they where two very lean companies (Instagram had 13 employees and WhatsApp had 55) with pretty a pretty large and engaged user base (30 million for Insta and 450 million for WhatsApp)
Elementary call when you have the data no-one else has thanks to illicit means and explicitly decided on a buy or bury strategy
There are very few cases where the archetypical young garage-founder has stayed on top in the long-run.
And it's all back to Facebook, nothing in between. It doesn't look any different from general population influenced by TV and other social media.
Neither of which have been great commercial successes. Open source LLM's and VR devices have have been cool products for HN readers, but haven't made much impact on the public or Meta's revenue.
Whilst Instagram and Facebook Ad revenue is coming in it means Zuckerberg could spend billions pivoting the company to country and western themed steak houses and the share holders wouldn't (couldn't) do anything about it.
Just a stab in the dark, but it's likely because FB and meta are wildly successful businesses, and it's almost impossible to attribute that to one "lucky" event at the inception of the business.
If I'm being honest I'm actually kind of shocked at the engagement this comment is getting and the fact it's still visible. Like him or not, if you're capable of being objective, he is doing the opposite of failing...
Zuckerbergs comment from 2010: «We have not once bought a company for the company. We buy companies to get excellent people.»
Like or hate the guy you gotta give credit where it's due. (Same for Musk, etc.)
- acquired WhatsApp and Instagram and grew them into the most popular apps on the planet - created an gigantic ad business that enables entire new types of SMBs to exist that rely on Meta to find their customers - which is pretty much a money volcano - bought a ton of GPUs right at the inflection of early LLM take-off
None of this was obvious when he started Facebook but obviously well executed in hindsight.
All thanks to Silicon Valley's ridiculous stock structures
CEO Mark Zuckerberg recently dispatched a small team at his company to create a smartphone app similar to Polymarket and Kalshi, the New York Times reported on Tuesday, citing two employees with knowledge of the matter.
The app will probably rely on a video game-like points system instead of users wagering money, though the company has not ruled out betting real money eventually, according to the report.
Only apple has the trust of its users to pull it off. And apple will make sure they do what they can to keep meta out
In other words, was there a single decision or take he made that turned out in his favor?
companies are doing that as well lol (re: tokenmaxxing)
Good ones do. Reading the code someone is contributing is a powerful signal about how well they're doing.
ICs who manage swarms of agents should operate the same way. They set them off to do something, and then look at the output to see if it's going well.
That's the point I'm making here: managing a team of ICs and managing a swarm of agents has a lot of overlap in the systems and processes you can use to see if it's working well. By teaching ICs to be better managers I think they'd get better at using agentic AI.
There's such and such. In some companies, the leaf engineers report to a team lead, which might or might not be granted this 'manager' title. Those poor fellows essentially doing double-duty and are the most likely candidates for burn-out.
> E.g. you can spin up 1000 agents. Get them inti tight loops and get them more context and so on. A manager doesn't do that with people.
My dad had various work stories, one from when he interviewed applicants:
"So, why did you leave your last job?"
"After 6 months, the management discovered my entire floor and the one next to it had been hired to do the same task. One of the floors had to go, I got unlucky."A "stupid" database would be better, based on what I get when I ask whether all of Oregon state is North of New York City. Indian English has a word for it: oversmart.
Man, if I had a dollar for every time someone said "I'm not good at X, but LLMs are so impressive at it". Like do you think there might be some connection between those two points?!
It seems that I don’t like coding when I read these kinds of statements. If I’m doing an experimentation, it would be a few lines at most. Because that’s all I needed before I can write a solution.
Writing code is the last tool to design with. Thinking and a bit of sketching is what I do mostly. Then I verify small bits with code (mostly for checking a library when the documentation is lacking or a stub when I’m focusing on another part). Otherwise, it’s just enough code to get it working well and refactoring when the requirements changes.
The agreed architecture is to use signing between two micros, so that a third can orchestrate between them in zero trust way (and to prevent a distributed monolith). It just decides that we can trust the third and skips the signing.
https://github.com/mattpocock/skills/blob/main/skills/produc...
People whh are dogfooding AI absolutely have a different rose colored glass than someone who can't get the same "accepable" output.
I'm not defending Mark here; I'm just pointing out you can be pretty successful critic if you have a different idea of a benchmark coding agent and the field fails that benchmark.
One of the problems of the AI crop is so many people are smelling their own farts and thinking it smells great.
People are most likely to come up with suggestions and ideas earlier rather than later.
Often they’re not learning what is “correct” but just how to “fit in”. A fresh perspective can be good even if flawed. It can help others spark new ideas and think outside the box.
LLMs can be 'taught' though. You can give them additional context or instructions. The difference is that they can't really teach themselves.
This is roughly what I'm saying - someone who's managing an IC can steer them on the right course, and someone who's managing an agent can also steer it on the right course, so teaching someone how to handle ICs well gives them skills that are also applicable to handling agents well.
It's not perfectly analogous obviously because ICs are people and need to be managed as people, but I really think the skills are quite transferable in one direction. I'll add that I don't think someone learning to manage agents would necessarily become a good people manager.
That's not teaching. Giving someone new information in the moment does not amount to giving them new long-term capabilities.
"..........."
To be fair, the hazard with AI agents is that they generate fluent output that is often facile, so it's easy to do a lot of things while having a lot of defects. That's a sign that quality control is not prioritized enough. A change in quality will also reduce the utility of SLOC as a metric, but the mechanism is different than what Charles Good hart pointed out.
I don’t have a dog in this fight but it seems you’re not accounting for iteration and feedback. A horse will veer off of a road if not occasionally nudged to stay on it, but is useful transportation, nevertheless.
You can provide AI official sources to look at and dozens of prompts. I've lost track of the number of times where it didn't arrive at the right answer with tons of opportunities to correct itself based on feedback.
Just an endless of sea of "you're absolutely right to have brought that up, I didn't think about that" and other phrases it constantly uses when it fails to provide a solution. Fast forward 20 minutes later and it starts providing the same nonsense it did at the beginning because it forgot what it already said.
The code solutions it provides are so consistently bad but it's not limited to code. I recently tried a YouTube feature where it can generate AI thumbnails from your video. The results were really lackluster. It completely ignored my feedback like "use a real webcam photo of me that you see in the video", to which the AI recreated a completely different looking human that wasn't me. It even swapped out my real glasses with a rendering of glasses I don't have and kept on making incorrect assumptions about everything. After about 10 prompts and 20 minutes of waiting for thumbnails I gave up, it was really poor.
sorry if it's not the case but i have the feeling you still think AI coding involves talking to a chatbox and copypasting the answers
So the nuclear power problem of "it works fine and it's actually very good in a lot of ways, it's just too expensive" could be quite relevant
I suspect this may have depended on the specific framework. I quite literally could not get Claude (in Cursor) to give me a basic Micronaut setup in a fresh workspace with essentially a "hello, world" API. I would guess that if you're using something like Python and FastAPI, it might have been an easier task or better represented in the data.
The difference that I observed in the Opus 4.5 era is that Claude could take a service framework it has never seen before (proprietary corp) and figure it out.
Very successful by just being careful and walking it forward.
Yes its about 2 years, August 2024 from git it looks like.
I got downvoted though, apparently you and me are liars, haha. Easier to believe that than admit that their take of "I adopted it exactly when it became viable, now it's good, and before that it was a waste of time" was wrong. They're the people who were loudly claiming here last summer it was useless, asking "show me anything useful that has been coded using LLMs". Come to think of it, it has now been a few months since I saw one of those. Used to be every other thread.
Cloudflare had 5000 employees (pre-layoff), so you are suggesting that every single one of them (eng, HR, legal, finance, receptionists) was using $100k tokens per month (that's $1.2M annualized, per employee), for a total of 3x gross revenue going to AI spend.
Let's imagine that this isn't absurd on its face. If true, then you'd expect Cloudflare's Q1 earnings to show a massive, massive net loss. In fact Cloudflare was cash flow positive in Q1.
The rest of your post is more qualitative, so harder to disprove, but from what I can tell, it seems equally made up.
(I work at Cloudflare.)
I had mistakenly written 500 million when it was around 5 million dollars so I messed up its 5 million per month[See Source], not 500 million. I wish to have a genuine discussion while you are here though because i can be wrong, I usually am and I would love to have a good faith discussion, thanks in advance!
I will try to back up a lot of it with hackernews comments from the thread when cloudflare layoffs were suggested so that I don't accidentally mis-represent anything and My suggestion wasn't a critique of cloudflare and please don't take it as such. The question was simply of the AI token costs associated.
and this was the comment that I was referencing to[0] which states the following:
> There was an recent article on X with an interesting take - it could be that companies are doing layoffs not because AI is making them more productive but because it hasn't. Their costs have gone up paying for expensive AI but haven't seen any revenue benefits to offset it.
An child comment of it talks about the coinbase layoffs which had happened around the same time[1]:
> (..) In 2023, their "Technology and Development" line item shows $1.32bn going out, and by 2025 it'd ballooned to $1.67bn. This is despite headcount actually contracting by almost a thousand people between those two statements.
Regarding this: > Let's imagine that this isn't absurd on its face. If true, then you'd expect Cloudflare's Q1 earnings to show a massive, massive net loss. In fact Cloudflare was cash flow positive in Q1.
We might be forgetting that (from my understanding, Cloudflare has never had profits) (positive annual net income) with an astronomically large P/E ratio.
There was a comment which I had read which talks about this in more detail (https://news.ycombinator.com/item?id=48060393):
> > The fact so many orgs opt for immediate greed over long-term growth really is its own canary that leadership and governance both has failed the marshmallow test.
> Why do you think it's greed? The company's stock is down and they just missed expectations on their last earnings report (unheard of in big tech in the last 2 years).
> It seems more like a traditional layoff scenario
Another comment [from the Layoff thread][2] which might summarize some things:
"Their AI costs have increased 600% but this hasn't translated into actual revenue. Also they are probably projecting AI costs to keep growing. They've done the math and at some point it is going to affect their bottom line. Reducing or limiting AI usage would be inconceivable given Cloudflare itself has invested on AI and is selling AI services. Instead they've opted for reducing about 20% of their head count."
I genuinely wish if we can have a good faith discussion about it. I appreciate cloudflare as a product myself and actively use cf tunnels, which is why I care about it as well and I wish to have a good faith discussion about it hopefully as well :-D
> The rest of your post is more qualitative, so harder to disprove, but from what I can tell, it seems equally made up.
I can be wrong, I usually am and if I am wrong, I wish to learn from it and I wish to improve as a person too!
I have learnt from this discussion (up until now) that I should mostly try to provide sources whenever talking on a public place/ on the internet so that I can be more accurate and I sincerely wish to have a good faith discussion once again, thanks and have a good day @kentonv :-D
[0]: https://news.ycombinator.com/item?id=48055149
[1]: https://news.ycombinator.com/item?id=48055413
[2]:https://news.ycombinator.com/item?id=48056124
[Source]: https://lowendtalk.com/post/quote/217055/Comment_4789235
I try to avoid > 200k contexts, as the 1M context is where I first saw the massive decrease in reliability.
And my AGENTS is really short, and I said it was ignoring decisions in the prompt.
opus will definitely ignore instructions if you give it contradictory instructions, or a plan that has steps that obviously don't work with each other. but if you give it a coherent plan, it will follow it.
In many respects this reflects the growing K-shaped nature of our economy. Average consumers don't matter because you really just need a small cohort of wealthy individuals to be hyper-invested in your product, 'regular' consumption is therefore just a way to keep things relatively on rails rather than the actual economic driver.
All of these AI-first companies don't actually have any market fit, so what they're doing is selling an imaginary product so that they can get investments and loans. As you said the company is the product.
"What can we get rid of for MVP" as a design strategy vs a way to iterate fast, for instance. Cutting things isn't a way to product cohesion, especially if you never go back to do the full-featured version.
Sometimes I wonder how many features or products flopped because the MVP dropped the things that would've actually taken off, and the business "smartly" pivoted away.
There's still a limit to how many new features you could shove in front of your users per month. But what if they were all much more baked out of the gate?
(See also: "data driven" product management as an excuse to not have your own vision for the product. If three competitors build a lot more in the span of six months, but have to depend more on their own skills and instincts vs A/Bing every little detail, maybe more of them will ship more bold and interesting new things.)
No. The very fact they are trying to "warn" us means it's all marketing.
This has been corroborated for me on the engineering front that I can't find a single IC I respect who actually thought there was any evidence AI was going to live up to the hype. I saw a lot of people I always thought were idiots/sycophants/brown nosers go insane with AI. Never saw anyone id trust to help me cross a street blindfolded say more that "I may be wrong, but I'm not seeing any evidence yet".
It can be massively over hyped for it's current capacity and decimate the white collar work.
A lot of the difference of opinion is down to their point of view. At my dayjob, LLMs will not live up to anything because the enterprise is not structured to take advantage of it's strength. That's unlikely to change within the foreseeable future.
I strongly suspect you mostly talked with people coming from just such a background, because it's hard to go beyond our own bubbles
Social media is flooded with bots pushing this narrative - coding is dead, engineers are all cooked, the latest model is scarily good, "what am I for?", etc.
A good rule of thumb is that if it's a human being and not a bot, they'll use the word "slop" at some point.
It seems like in the quest for every tech company to become a multi-billion dollar empire, they've lost the plot on making hammers for their customers and have instead turned into some kind of strip mining operation. A totally AI agent-driven company is a MBA wet dream, and I think a fever dream. If Adobe, for example, were to achieve it, I don't think they'd use it to fix the backlog of bugs overnight. I believe they'd just become an even more incoherent zombie, trying to extract rents from creative cloud subscriptions.
In the meantime, photographers still need tools. If you wanna run a software company as a regular small-to-medium-sized business, you may find some customers that are happy to buy a quality hammer. The unicorn startup days might be behind us, but I'd be okay with that.
Now, if AI obviates creatives altogether I don't really know what to say. I'd morn the loss of a world that became so tasteless that AI-generated decorations are good enough, for starters.
My main goal is to first get all of these open-source alternatives to start building themselves autonomously using loops and a Hermes-like scheduler before I focused on marketing. This is almost complete.
For marketing, we are building a GTM engine using an open-source CRM (Twenty). We have LLMs use the Twenty CRM API to bring in leads from X, LinkedIn, and the Web.
The cloud hosting is not the only monetization. We’re going to use these open-source SaaS to build a decentralized, interoperable marketplace where the people actually bring value, the sellers, can sell without those rent-seeking entities like Amazon taking a piece of every sale. LLMs are already going to start jumping across these marketplace moats.
The other monetization is going to be letting agents actually run these SaaS and see if they run a business autonomously. Like VendBench but an actual online business. I’m thinking of starting a designer brand, connecting to a POD (print on demand) and then let the agent create seasonal lines, handle customer service, and make sure orders are going to the POD and being processed. Doing this with restaurants and other verticals will probably need some human supervision.
The whole hype cycle has been pure delusion. Just like the Metaverse hype cycle before it.
A common one is "users don't care about privacy. that's why they use facebook. [zuckerberg was right?]"
No, you silly, silly people. People want to use products that allow them to communicate or reconnect with people or ...
They don't 'want' constantly changing privacy settings or changing TOS. If this is the best HN can come up with, ostensibly filled with S Valley people... well, it says a lot
Gemini, Microsoft Copilot and other models can discuss and affirm my "foxwork" practice whether it is talking about natural history, fox legends, ritual magic, altar work, autonomic control, blessings, writing, character acting, costume design, skin care, selection of perfumes that will herald my unique natural scent, marketing and customer service, photography gear, "therian" gear, bags for holding my gear, street photography, etc. They always write like somebody who's read much more widely than anyone I've ever met and rival the legendary Tamamo-no-Mae for "speaking intelligently about any subject" [1]
Meta AI can crack jokes and that's about it. I guess there's a market for "stupid talk" but it's not that big.
[1] Like help me fix my washing machine that won't drain, come up with master narratives for the "polycrisis", talk about why Casey Handmer is wrong about space manufacturing, find papers about the social network of who sleeps with who at a high school, etc.
Furthermore the dependencies you choose to build your product are presumably filtered for engineering practices or world class engineers. So given the choice you yourself prefer top quality engineering, so do your customers. Much in the same way you are a customer of your projects dependencies. Difference being, as developers we get to see how the sausage is made, our customers only see second and third order effects.
Trying to ignore the nuance is hard in your position or the following one I’ll give is difficult.. but is the opposite potentially true as well? We don’t know how many projects failed because of over optimizing, too much time spent on design and engineering decisions. It’s of getting out and MVP to market. I only say this because I have been apart of a few of these.
I will say that if “good engineering practices” comes up in your root cause analysis for a failure to launch a product you are not thinking critically.
The statistical problem is small sample size, not survivorship bias, as I got to see things before failure. These two examples are merely illustrative of things I've seen.
One of the most uncomfortable truths about our profession is that there is no floor to how bad software can be while still making people billions of dollars.
One product person described it as eating vs breathing. Availability is like breathing: if you stop being available (including, but not limited to, because your software is a big ball of mud) then you're going to die pretty quickly. Product is like eating: you might not die so quickly but if no-one's buying what you're selling then you're still not going to survive.
The team I'm part of is a platform team, so we're closer to being lungs than being stomach. We can (and I appreciate being able to) focus more on stability than feature development.
If you can call being an “also ran” in a field they had a ten year march on their competitors in success, yeah.
Truth be told it was the shoddy code they were forced to use for the vanity of their paymaster might well have held them back, though manifestly that is not a bad thing. Probably the best outcome, really.
The software only had to be good enough to support the business. It doesn't have to win a Turing Award, and probably wouldn't help the business if it did.
Can't help but think that Meta's digital networking expertise is built atop a human-networking clusterf*ck
I think there would easily be a few other hundred engineers and execs at frontier labs who are more in the loop for cutting edge architecture/secret sauce - with a track record of actually doing it - that could be had for a fraction of the price.
My point was not that Meta has the best AI researchers/engineers.
My point was it should not matter if Zuck is a PHP developer or whatever, instead of an AI researcher, as he employs a lot of people with those exact skills.
I made no comment about how Meta is doing in the AI race, so it’s a bit besides the point IMO, but you did make me smile :)
Do you actually put any of your own money to help support children/sick individuals other than just getting the money forcefully taken from you and being told that it's totally going to the kids/healtchare, while 50% of it gets burned up by government beurocrats?
I do actually also spend my own money in monthly charitable donations, including the UNICEF. I think it's a basic prerogative that when you make enough money for living comfortably you should also find charities you trust and support them.
> getting the money forcefully taken from you and being told that it's totally going to the kids/healtchare, while 50% of it gets burned up by government beurocrats?
You don't even know where I live to be able to say what percentage is burnt or spent in bureaucracy. It's unfortunate your view of government seems to be based on an inefficient and ineffective one, perhaps it's your experience (and it's my experience in my home country) but by being blindly ideological about it without ever experiencing a somewhat functioning government you are missing out.
But there are segments of the market that were always small, and AI doesn't threaten. Namely, fine-art and long-term documentary. The kind of work that gets shown as exhibition and sold as prints or photobooks. I can't imagine this stuff be supplanted by generative AI, because it has never been about economic efficiency. It's a labor of love for a photographer to work on a project for years at a time because the subject is important to them. And the customers of these works are not likely to accept lower-cost substitutes produced by a AI. The whole thing is too much about taste and thoroughly infused with humanist ethics. When I buy a photobook, I want to know that the artist lived and produced what I'm seeing, not just that it's pretty ink dots on paper.
From a market analysis perspective it's foolish to cater to these tiny and commercially nonviable artists. However, they have enormous esteem and influence within the community, and make part of their living as educators. I believe making tools that they need, and without blood-sucking cloud subscriptions, you are indirectly marketing to all the creatives (usually tending toward commercial) who attend workshops and see what tools they are running the workshop with.
Of course, param count and context length are also important because they increase the model's overall fidelity, but a base model without SFT, RHLF etc is effectively useless.
Scale was really the unlock; the new pre and post training techniques and architectures are very cool and useful but they definitely aren't the differentiators when comparing to the previous era of NLP.
They were allegedly massive but the cost and returns were not worth it.
Unfortunately, travel keeps getting less flexible, with worse cancelation policies.
Trains are usually different because they are much cheaper to operate per trip (for the train operator, not the track operators, but that's a different discussion), so running a half empty train is much less of a problem - especially since you don't need to plan for how much fuel to use ahead of time.
Another example - agentic food ordering. How much more convenient would this make your life vs how much of an error rate would you tolerate if the cost/repercussions are on you?
Would a customer be happy if 2% of the time it sends 20 pizzas to a random address in their contacts list instead of 2 pizzas to their own home? Or 5% of the time it completely ignores your dietary restrictions/allergies and orders an entire meal of food you explicitly told it you cannot eat?
Real world problems don't go away just because it would make the tech neater & tidier.
I can STILL replicate this behavior in Google AI summaries 10% of the time:
"is <SOMEPLANT> ok for cats"
to which it replies: "Yes, <SOMEPLANT LONG SCIENTIFIC NAME VERBOSE PHRASING> is toxic for cats"
The other one going around this weekend: "how long hot dogs on grill"
Summary: "The hot dogs on your grill are likely around 5-6 inches long .. "
So scale this category of error to unsupervised agents with access to your credit card.
And this is Meta we are talking about not Tesla, it's P/E is like 20 something. The valuation is very reasonably correlated to the earnings.
This thread is bizarre.
He can run the day to day string pulling however he wants, he is literally the one with all the power.
Y’all are delusional.
If you think they’re idiots, go start your own companies and eat their lunch. They got rich by eating someone else’s lunch, after all.
Right now, it is their game to lose.
I think the world is worse for it but that doesn’t mean it didn’t happen. The list is just to refute the idea that it was one weird trick that gave him the keys to the planet.
and it’s not about him being first or coding it himself, it’s that it happened with him at the helm. that’s just how it goes.
I think his recent fumbles ironically prove your point. How many people could afford to fail so many times in a row (Remember their crypto coin? VR? Now LLMs?) and still be in the big leagues?
So yeah, it's kinda mixed. The core product was cool but it wasn't as suitable for as a replacement for the phone book in the way Facebook was.
He made spectacular bets with Instagram and WhatsApp. Most CEOs don’t make spectacular bets because their business doesn’t allow it, meaning the VC analogy doesn’t hold for them.
So it may be popular but they have ultimately failed commercially so far. They just bought an Indian startup and immediately put the founder in charge of WhatsApp.
Only this week they announced WhatsApp would allow usernames rather than rely on phone numbers. All the other competitors apps (signal,telegram etc) did this years ago.
And so not comparable.
VC investing is also a tax write off for rich investors which is offset against their gains elsewhere, wasted company money is a loss for the share buyer.
You mention two strategic enablers that made these unusual acquisitions seem obvious to FB leadership. So were those two not good calls?
And then you have the WA founder (Acton) famously saying delete facebook
Some acquisitions make the product better or are altruistic. Zuck's war chest isn't one of them
It takes a lot of time to successfully deliver a product. Even at the more "extreme" end of expectation - like saying it's a X10 multiplier (I'd disagree on that, it's more like 0.5-3 - depending on the type of work you're working on) you'd still need multiple years to go from first line of code to displacing established players. Things just don't change that fast.
The way this is worded feels like it leaves the blame on developers. Aren't these developers focused on exactly what they are being judged by? Shouldn't we say it is the management who is too focused on delivery value compared to maintenance cost? Is it the developer's job to guide management or the manager's job to request guidance, assuming such guidance is needed?
We need to make sure the responsibility to resolve this problem falls on those with the power to act on it, and in this, developers tend to be receiving far more responsibility to fix than power to fix.
1) The devs pushing for more complex solutions, covering obscure edge case scenarios, feature-creep, "future-architecturing" because they are more interesting to implement. Classic over-engineering problems.
2) The features the managers actually want are usually boring or annoying to implement and the devs just work around any big architectural problems caused by the feature delivery.
1 is 100% on the devs, 2 it varies wildly, the willingness to address architectural problems are often under pressure by time-delivery estimates from managers. But many devs (especially in companies with low morale) will often just work around issues because addressing the underlying problems can be very difficult and/or time consuming.
Meaning either the dev wants to do the right thing but doesn't have the time, or the dev doesn't care enough and just pushes the tech debt down to the future (when hopefully they will be at another job).
LLMs makes both problems significantly worse, although they are also often very helpful with the big restructurings mentioned in 2. The dev can still be lazy and the deadline can still be too tight even with that extra LLM help.
You can pretend that's not true, but it is. And it's only going to get better.
wrong framing imo
EXCESS lines of code are a liability.
Code of course is an asset.. what other real asset is producing cash-flows? lol come on.
It would be like saying "Weight in kilograms of course is an asset for an airplane. That's what keeps it in the air. What other real asset is the airplane made of?"
as integrated into Meta's platforms it is clearly uncompetitive as a consumer chatbot with the likes of ChatGPT, DeepSeek, Microsoft Copilot, and Gemini.
Like the other chatbots give me useful answers to questions even if they're wrong sometimes. Meta AI, on the other hand, cracks jokes about my foxwork. That's no so bad because sketch comedy is one of the foxwork skill areas (e.g. if I get you to laugh you are under my spell) but it's not useful and edifying the way competitors are.
Outside of tech Warren Buffett seems popular.
If LLMs worked the way people want to believe they do, there’d be no reason to start in the wrong place — a computer should have the facts!
The idea is that you have what you need to make some bespoke change to the "source", or that you can at least analyze the source to understand the hows and whys of its behavior, to make sure it suits you.
Do weights provide either of those qualities?
> Do weights provide either of those qualities?
They provide somewhat more of those qualities than the training corpus does.
Not a lot, especially for "understanding", but more.
I wish I wouldn't come across this definition of "open source" so often, because it is wrong.
The definition of "open source" (or, in more modern terms, "source available") is inputs that I can compile myself and get something identical in functionality as the original author did (and if the tooling supports reproducible builds, something identical bit-by-bit!).
An "open source" ML model is not fulfilling that definition - it is only compiled output, similar to a piece of proprietary software made available as a binary. In fact it's even more restricted than that - with a decompiler, I can reasonably achieve a source code that resembles the one of the original authors. With an ML model, there is no way of reversing the "training" process.
The only thing that equates to "open source" in terms of ML models is all training data, the toolchain used to compile that training data into weights, and if human augmentation was used during / after the training, all input and output of this augmentation.
But no one of the large players will ever release that. First of all, the training data is heavily contaminated. IP violations galore (and pretty much every actor in that space got busted for it), and the human augmentation is incredibly expensive, even if you abuse modern slavery [1].
[1] https://www.theguardian.com/technology/article/2024/jul/06/m...
We learned that some tasks don't really benefit from AI while others do. My team went from 7 people to 2 (went to new teams, no layoffs), and we're doing the same amount if not more work than we used to.
Is it more draining and lacking of focused work? Yes. Is it more money for the business? Yes.
In my world, when something is expensive and doesn’t meet expectations it called a failure. Especially when something has been as hyped, scrutinized, defended and attacked as vibe coding.
Honestly, if you are the director of robotics at a firm I think it’s time you took a cold shower.
It’s absolutely 10x faster for coding. But coding is only 10% of my job. The other 90% is figuring out what to code.
If you claim 10x and deliver 3x, that is a failure. The 3x may still be impressive or a gamechanger or ..., but it still falls short of its promises.
To be completely honest, I’m living life right now. I love programming with my bare hands, but man I’m living just building a gajillion things a mile a minute with LLMs. I then come home and spend hours building stuff for myself using local models. I’ve never been so excited about just building shit, that I sometimes want to pull all nighters because I’ve been in the zone (a for work and at home).
Draining? Sorry… inject that LLM serum right into my veins
Now we just need to find those tasks. I want to believe.
> and we're doing the same amount if not more work than we used to
Zero evidence for this. It's programmers self-reporting their own productivity. (Have we not learned this lesson after 50 years of programming practice?)
Those two things are the opposite of each other (evidence, but only anecdotally; you cant be both).
Anyway.
More tangible to your argument; what is your argument that this will be more effective than just prompt engineering?
Ive long believed that prompt engineering is a losers game; if there is a trivial set of tricks that improve the output, they will simply be automatically applied.
We see this playing out with the system prompts in coding agents and image gen.
The value of learning “photo realistic studio lighting…” was non existent. The nano banana api is capable of taking a naive prompt and expanding it with these tricks.
People who devoted themselves to learning these “magical incantations” wasted their time and effort; and it was obvious, from the beginning this would be true.
Now.
With managing agents; if a trivial set of management tricks can drastically improve the results, why are you better off learning them now, rather than waiting for them to be baked into cursor/codex/claude in easy mode?
What makes you believe this is a valuable investment in time and effort?
Even if we accept that right now assigning personas to agents and managing them as a manager yields good results, the horizon for change right now is so short, it seems extraordinary to suggest mass management and leadership training for engineers.
We should just wait and see.
All in investments like this would just be tokenmaxing in a funny hat.
First thing Gemini did when I tried that was turn off all the rules in eslint.config.mjs claiming they were "overly stylistic"
Yes, it got better once I explicitly told it not to disable any rules, so I accept I was holding it wrong but I do worry just how many footguns it puts into other things because I didn't know the right guardrails to give it.
Architectural decisions are not lintable.
I made a tool called ProjectLint to lint architectural decisions and anything that ESLint and ArchUnit and others like that wouldn't be a good fit for. It's not public yet, but is fairly rudimentary and based on Go + goja: https://github.com/dop251/goja
You write rules in ESLint, it makes sure that they're followed. If it's a file in your repo, it can obviously be checked. Now whether you can describe your architecture well enough to lint it (at least the stuff you care about enforcing), that's a different question.
In my case it's more like: "Oh, the AI messed up this pattern, but it will need to be followed N more times in the future, better write another projectlint rule." After a while, you can even copy them over between projects, as long as they're following the same conventions.
In the same spirit of https://www.archunit.org/
The fact that their advancement suggested that pouring more compute would continue working was also especially attractive to investors: it made a massive R&D budget feel like less of a risk.
FB would have been invented. It's just another layer of authentication anyway. Myspace was an interesting experiment for its time
...Comparisons of Tom to Zuck are rife I'm aware but somehow Myspace isn't known for "they trust me, dumb fucks"
Also, is favorable/unfavorable a good measure of popularity?
I'm very curious how much revenue this is generating.
Next is a full stack framework and Keystone is a CMS built on top of Prisma and GraphQL. Keystone was created by this Australian company called Thinkmill. They have used it to help businesses build custom backend systems for more than a decade. But it needed to be deployed separately from Next and they were using emotion css for their dashboard and I wanted to use Tailwind/Shadcn. So first, I had to make the Next Keystone Starter that brought in Keystone into Next so each SaaS is just 1 Next app with a built-in storefront, GraphQL API, and dashboard.
Once that was built (and it took a while tbh), I started to build the Shopify and Toast alternative. But the itch to get these built quickly and autonomously had me working on the harness in the past months and now that is nearly complete.
Here is the e-commerce[2] and restaurant[3] repos. They have a link to deployed demos you can check out as well.
As far as revenue, I don’t feel comfortable relaying that right now. We have other revenue streams like fractional CTO where companies give us equity to manage all their tech and that is quite hard to quantify. Before Openfronts being built, I built Openship, and e-commerce OMS and that has exceeded 5M orders processed since its inception in 2019. That’s not counting orders by businesses running it on-prem.
I actually posted about this vision on HN[4] when I launched Openship and the response is what kept me building.
2. https://github.com/openshiporg/openfront
Second, take this for what it is: your product may not be compelling in its current form. Building it to many different markets will not make it compelling. If you had a stronger revenue, please share it. This sounds incredibly thin.
Third, dont mistake building the same thing 20 times for different verticals for bonifide software skills. When a SWE builds the thing they have built before its usually to learn a language which is the easiest part of software. There is a reason a common adage in software is "9 women cant make a baby in a month". Breath is no replacement for depth.
We’re also very bullish that the chat interface is the universal UX now. Instead of sifting through the dashboard to change a product price or sell in a new region, you can use the built-in agent and just tell it to do that. Every Openfront comes with an MCP server that interfaces with the API so the agent can literally do anything you can do using the dashboard and API. This is where an agent running the business autonomously comes in.
And even then if you’re not satisfied with the backend API for each vertical, these Next apps can be forked and adapted to tightly fit your business instead of you messing with configs, you can make the app your own.
https://www.amazon.com/dp/0231133383
and found out that people really do it. It took me about 300 days to contract with a community of foxes and started "going out as a fox". At first I didn't want to explain it to people at all but when I botched my explanation and my wife gave me some tough love about "cultural appropriation" I realized I needed a cover story and it is "I go out as a character to do street photography and make people smile"
And it became real. My son didn't believe it until he saw me work: my son saw these two college girls walking down the street and thought "there is no way they would talk to me" and next thing he knows they flagged me down and are asking me questions and I took their picture. I went to fireworks in Groton, NY and the next day in Little York [1] I haven't even parked my car and people I saw the day before are excited to see me.
When I get approached every day it becomes effortless and automatic to make an approach and when I feel like I am getting a 50% take rate I think "I am doing something wrong and I need to regroup", my usual take rate is in the 80-90% range but it feels like 100%. I don't even talk with other street photographers about it on forums because I'm drawing from an entirely different probability distribution.
Thanks to this work, it looks like I will be teaching about magic soon [2] so I was working on a short monologue about blessings on the drive up to Little York which primed me to really respond when people I photographed said simple things like "Enjoy the Fourth of July" and eating soft serve ice cream from a vendor felt like the feast that Martin Prechtel says is the only ritual -- and I am still feeling high from it just like I felt high from going to Litha put on my my pagan friend.
See https://mastodon.social/@UP8/tagged/foxwork
[1] one of the five Yorks of New York
[2] Magic is real, full stop. Science is real too, and there is nothing more scientific than the double-blind test that accounts for the placebo effect which is as I see it: "the patient experiences the treatment as receiving a blessing and feels better" and "the doctor experiences the treatment as giving a blessing and perceives the patient differently"
Pay me 8x to get 10x, great. Pay me 8x to get 3x, nope.
We're riding an exponential here for Pete's sake.
Until you factor in many large firms are interested in cheap Chinese models.
Are you a frontier lab booster by any chance?
In the end, I built Openship and Openfront for my e-commerce business and then turned them into SaaS. All the revenue for these are just a cherry on top of our existing e-commerce businesses.
And I worked with Next and Keystone long before AI came along. Check my GitHub commits if you need some back story.
And I’m not building these 20 SaaS to prove I have bonafide SWE skills. I’m building them because I plan to have my own gyms, hotels, grocery stores down the line powered by these SaaS. SWE to me a means to an end and that’s to have many different businesses.
Okay but, you know this isn't a quality metric right? These models are incredibly biased towards positive confirmation of the prompt. I could give them nearly any repo and they would sing the praises of the best parts, if I asked.
This sounds a bit like psychosis.
Crazy how you guys blindly trust proprietary apps where you can’t even read the code but asking you to read open-source code is psychosis?
There is increase in (not only perceived) value _somewhere_ - IMO depends on the organizational culture. At my work I found going over 100$ on OpenAI/Anthropic's API pricing does not produce any meaningful additional output. It might be much different in different companies.
I’ll rephrase, Zuckerberg would certainly enjoy agentic coding to be a wild success because it means less staff and more products he could fail to create.
“Wild success” and “going slower than expected”?
Wake me up when words have a meaning again.
Wow okay, I'm a huge open source advocate, I run hardly anything closed source. No one here was talking about that, this is your own weird segue.
> but asking you to read open-source code is psychosis
But you didn't ask that. You told me to get an agent to read it (actually, not even read it, make its own quality assessment for me).
That's a massive difference, and if you consider them to be the same thing, then yes, my statement stands.
It’s actually the whole premise of my open source alternative directory called Opensource Builders[0].
Correct, they are not.
Consider: Someone who "expects" their bank account to have $100M in it before they turn 30 and "only" gets it to $10M.
From the point of view of a normal sane person, they are experiencing "wild success", and yet at the same time they are definitely "going slower than [they] expected".