GPT‑6 and Intelligent UI for everyone(openai.com) |
GPT‑6 and Intelligent UI for everyone(openai.com) |
Ok so this is how you build a moat. Models are not a commodity and visualization isnt a purely harness problem.
> We expanded our training methods to help the model make thoughtful decisions about content, layout, visuals, and interaction. This included evaluating the interfaces it creates for clarity, usefulness, and completeness. GPT‑6 learned to use the component library and make good design decisions, including how to organize information clearly, when to use interactivity, and when a simple text response is enough.
Curious to know how they trained it produce appropriate visuals. And why that cant be done with propmt engineering.
Or, more probably IMHO, most people will keep using the standard UIs because they are ready and they need no work to build.
Now they extracted the value from interoperable open systems, they would love to replace it with closed proprietary systems.
How do sites like Wikipedia get updated? What if I need to upload important documents to different portals? It all has to go through the AI companies? They have access to medical records too? My taxes, social security, etc?
Is this how Microsoft imagines Office 365 and Sharepoint will be merged into? Disney, NY Times, Youtube, Netflix will all be fine with chatbots handling their content?
yes you and everyone else.
This is false.
Obviously if your model is stuck inside a CLI terminal, then not so much. But in a GUI harness (shameless plug for my own one: https://juggler.studio, but I assume others can do this too), you just ask them to answer in HTML and they'll happily draw pretty pictures inline in the conversation. I've been doing this for ages with claude, GPT, Deepseek and others.
Unsure if this has been done already but the best thing I did was get it to create a list of next steps as quick buttons that it keeps updated in a docked card. It works well but occasionally gets stuck on some un related rabbit hole.
Supposedly, people were struggling to follow written manuals and they needed an interactive explanations.
I'm not a mathematician and I would certainly love having a tool that would do ELI5 on some complex stuff, but I'm really worrying about using this too often and outsourcing my ability to do stuff to some mega corp.
This is now trying to steal YouTube repair and cooking videos. Here is news for you: People prefer YouTube repair and cooking videos.
But getting actual service manuals cheap/free these days can be tricky, and their quality can leave a good bit to be desired.
"Here is a library of [svelte/react/whatever] components, use them to construct a helpful visual to demonstrate your point."
The deconstructed bike at the beginning was in a class of its own, however.
I guess the above is mostly mobile focused.
What is it supposed to demonstrate? That the model knows some kind of folk mereology?
I suspect it will struggle with anything that isn't simple enough for demo-mode or lifestyle fluff.
Remember when Steve said ‘the computer for the rest of us?’ We’re seeing that here.
why not just stream html?
For elements that manage layout and drawing things like these explainers, though, it's hard for me to imagine how that would work. My immediate thought would be to an abstraction over it, which is what it sounds like they did.
What's funny about "Intelligent UI" is I said like 2 or more years ago, that these AI companies need to start thinking outside of these basic chat UIs, they do some things here and there, but its really depressing how little they do to innovate in these spaces. Same with the coding harnesses, the UI for all these things could be drastically superior.
Fairly likely: https://medium.com/the-engineering-brief/openai-hired-400-ap...
But this is just a marketing page for some tech company, right? If only it stopped there. I have to deal with this nonsense in news "articles" as well on occasion, when some web designer intern is allowed to larp as a journalist for a day.
sigh Just give me text to read.
This is why I use an IDE instead of a TUI agent... I'm on a computer with a multi-megapixel display, I want to use it. If I could be driving the same process with my Macintosh SE as a serial terminal, what the hell is the point of my recent Macbook?
I would prefer if the AI companies stuck to just creating better models and making them as cheap and accessible as possible. Let others build the products. I don't want 1-2 companies to own every product in the world.
That seems to be the plan...
So this is a taste of the slop that's going to invade everywhere in a few months.
Also reasonable: it's possible to hire a world-class ___ to do that job better than AI.
Anyone can having blog posts as good as this by using AI. Just not by zero-shotting it with a 2 line prompt. It still takes a lot of work.
Given that OpenAI is making noises about merging work with chat (a horrible idea imo), and Work is very similar to Codex... I dearly hope things like these won't have any meaningful cross-polination into the actual work tools.
Seeing that the chat is based on 6.0 and not 6.1 is disappointing. The "Visual and interactive explanations" seems genuinely useful, but 6.1 is just so much better. I wouldn't truly trust the 6.1 with the explanations, but I'd trust them a fair bit more than 6.0. I understand that compute isn't infinite, but tons of people only interact with the chat and having your "things-explainer" be as good as it can be is important when people increasingly treat AI models as the source of truth, or even use them for academic learning and whatnot.
Still, the models will improve, so the visual explainer seems pretty good as an idea / mvp.
In all seriousness, his lovingly and expertly crafted explainers are still going to age like a handcrafted heirloom clock in a world of plastic-clad quartz movements. But it’s absolutely incredible that we are now in an age where a computer can manufacture a serviceable interactive explainer on whatever niche topic you desire.
> So too, is his turn to be relegated to a relic of his time.
No, it'll be tasteful artifact, not a relic. Records, fountain pens and automatic watches did not die. They are used by people who discern things, and no, none of these things have to be expensive (i.e. Neither Seiko 5, nor Lamy Safari are expensive, yet they are as dependable as their 100x expensive brethren).
Human touch still has that finesse and warmth.
You call it serviceable, but what purpose does this service? What do you now know about 7-speed bicycles that you didn't know before?
I've been learning music lately and it kept re-pasting the same one chord visualization throughout many conversations, almost randomly and often barely related to the question. So I at least hope this won't be as aggressive so I can prompt it away!
- Audio volume has 2 levels: on and off
- Play button worked exactly once for me: it played and looped the video. Couldn't be stopped afterwards.
- The video progress bar has no visual indication of where it starts and where it ends..
Is this what Phind died for?
Even TV channels like CNN and FoxNews don't have a decent one.
It's like we need ASI to have a proper embedded video player. The ultimate software challenge.
Gettin closer to fully dynamic interfaces for a lot of software. Hell, give me a mode in Google Docs that takes every pixel of chrome away then vibe the rest as I need it. Persist across new documents going forward.
I have explicit instructions to tone it down. Less taglines, eyebrow text, subheadings, decorative spacing, pills, cards.
Hopefully this doesn't bleed into the chat...
Getting a UI that explains the parts of a bike is ok I guess, but isn't it simpler to get an actual breakdown? Google "parts of bicycle breakout" gets tons of useful images instantly.
Getting a specialized app to split a bill? It was already trivial to put in a calculator if we cared to go item by item on the bill. Having to provide names and tag every item as I go is just more work. Usually real people just go $total divide by 5, I had more, let me chip in an extra $10.
Same with booking travel and wedding plans, these aren't things people would even delegate to a trusted friend usually, much less a one-off request to an AI bot or custom UI.
Most of the useful tasks it can do right now are research, technical question/answer, coding. In terms of "build a flexible ui that solves real-world problem", if existing mobile app isn't useful in this arena, then its unlikely a completely custom UI will do the job.
That said, I've done a few small ones like a quick one to practice alphabet of a foreign language, or prototyping a web game, and the like. But ultimately its nothing that is worth trillions of dollars.
seems a bit like a snake eating its own tail
1) I want a quick answer, and I don't care for the boilerplate UI. For example, if I ask how to make pancakes, I make them all the time and just want to a quick reminder on the ratios, but it might trigger a full UI that I need to sort through to find information.
2) If I ask a to me unrelated to UI question and it triggers a big UI build that is completely off topic for my question (meaning I'm desperately pressing the stop button and prepping rewriting my query)
Or the other one I see, for example if I look up a unix command like:
"ls all hidden files in the /xxx directory"
And I get back:
"Sorry, I am am unable to find /xxx in my current environment"
- People don’t have to stick to a single provider. They can all claim the same 20%.
And now we are asking LLMs to solve it.
> Also pick up garlic, rosemary, thyme, lemons, honey, almonds, olive oil, gravy ingredients, mint sauce, crumble topping and vanilla ice cream.
what??? how the hell are you supposed to remember all that. Even the before example tells you what kind of potatoes to get. tbh it would be cool if you could just click a button and get an order pre-filled out on a grocery delivery/pickup service. I really wonder who reviewed this post and if they cook, because imagining yourself in that scenario and reading those instructions falls apart very fast
Kinda just feels like google search results in AI, which imo is a step down from distilled information. Chatgpt is already able to generate charts and visuals upon request.
Curious why that feels condescending? Like, the average person who is seeking a recipe needs a photo to know what to shoot for
Prepare to lose essentially unlimited chat mode.
I imagine many will move to claude, as I will (return), unless anthropic makes more blunders.
Random link I found looking for the twitter post
https://pasqualepillitteri.it/en/news/21024/openai-merge-cha...
The only things that OAI still have over A/ are better coding agent GUI, no 5hr limit on >$100, much more reasonable cybersec guardrails (that allow most RE work) and... that's it.
I should have been more precise than saying "just give me text", but it was what came in to mind when I wrote it, as I was thinking about what an article is meant to contain as its base element.
On the front page right now - https://news.ycombinator.com/item?id=49980626
OTOH, I still believe the exploded view on https://ciechanow.ski/mechanical-watch/ is something else.
For one, it has real physics on the weight, and second it always shows the correct/current time.
FWIW, his all animations has proper physics to begin with.
The entry you posted is nice, but Ciechanowski is still peerless.
Buckle up boys