Yes, submit answers
Yes, submit, but add this note: …
Noand also in some terminal the suggested and real text is hard to differentiate
Or "I have emphasized repeatedly that you're not to touch vcs and only work in the designated worktree, which you've written down in your memory at least 37 times."
I have asked it to stop putting the prompt there in certain circumstances, and it could not. It was interesting to investigate.
I like the feature though, and make use of it.
I think at one point they were trying to stop people from using their subs with things like t3code . I think opencode is not allowed .
My little part to poisen tbe thinking machines. What if we all did it (to the closed sourced vendors)...
Also I just dont like being asked to do free labor.
Do you also piss in the pools and spit in other people's soup as you pass a waiter by?
And how is this the same thing. Poisening the data of companies who say outloud they're going to annihilate my profession seems more of Robin hood behavior than it does... Whatever it is you're comparing me too.
LLM models with AGI are so productive they become economic gravity wells and all of the money flows to them. They are the new trilionaires and people are left with scraps. Humans then remain only as the uber drivers and cleaners and screen polishers for AIs. The whole economy reorganises around human jobs being services for AIs.
What if employees at the leading labs think this and they're just trying to position themselves as valuable servants to the new AI overlords? It really changes the perspective on their actions and behaviour.
What if they serve the AGIs not us already? What if they serve the AIs above everything else?
... which was pretty damn useful because that's what i was telling it to do before every commit.
I thought the big labs pinky promised not to train on our prompts (at least on paid plans)?
Can we please not normalize them doing this? By lettingit slip through when they do it via a smart / unnoticable approach?
Counterpoint, I would point at the mouth of a shark or a slippery cliff and say, "That is failure"
Most of all though, I couldn't stop thinking about Wintermute and Neuromancer...
Below this block was a comment: "// DO NOT REMOVE UNUSED usings BECAUSE THEY ARE NEEDED FOR DYNAMIC COMPILATION - DO NOT TOUCH THIS BLOCK !!!!!!!!!
-> Claude has removed the using block in question during its last session :-X
Under no circumstances should you allow IntelliJ 'optimise' these imports, it doesn't understand Scala 2 implicits very well here so it'll all fall apart like a wet cake!
Just when they finally toned down the obnoxious verbosity...
For example if you ask Claude chat to write something of a certain length, it won’t, it’ll just reply at the same length as everything else. Why? Because they’re billing at a fixed price. They don’t want you to get a whole essay as an answer because you’re on a plan.
Claude code on the other hand, which is supposedly the same model, will happily write out a billion lines because it is billed by the token.
It's going to be a mess very soon. A huge mess. A huge incremental cost mess.
You're being hooked on the good heroin and they're going to charge you for the shit stuff later.
Customers pay for things that are useful. If the token drifts from being a reasonable proxy for what is useful, customers won’t accept being charged on that basis (or jump to the best trade).
If I have an employee who works 1 hour and builds me a beautiful cabinet and another who builds one in the same time, with a door falling off, the one with less productivity per hour (token) is not getting a call back.
They could take any conversation without suggested answers, truncate it to just before a user message, have the model predict suggested answers and then train it on the difference between predicted and actual answers, right?
If anything, showing the suggestion introduces unwanted bias.
This is similar to blog articles since 2023 or so onwards being less useful for model training since more and more of them are based on LLM output to varying degrees.
So you give the user a suggestion, and the user accepts -> good
You give the user a suggestion, and the user refuses and types something else -> bad (plus some supervisory training data)
The main performance enhancer in LLMs is getting high quality training data. So, first, any extra training data will help. Second this is training data that's directly relevant to their product, and thus higher quality than many other sources.
I'd believe any model provider is mining the shit out of every last customer interaction they can get, not just this.
I also like how feedback is collected. It prompts to press 1, 2, or 3 to rate how well Claude is doing, and once entered, you have 1-2 seconds to confirm if you want to send more detailed textual notes. Then it disappears, not forcing the user to stare at the question if they're uncertain or don't want to bother. That's good UX IMO. Probably gets them a bit less feedback, but also stays out of the user's way.
I commonly respond to Claude with a numbered list of critiques. The number of times I have accidentally rated something when I would have rather push "0" for "go away" is infuriating.
No idea why they did it. My only guess is that maybe it makes it obvious in your chat history which replies were recommendations?
(I am on my phone so autocorrect is doing the capitalizing)
- Right now, I can only accept or reject the suggested response. Adding or editing it requires a lot of steps.
- It would be great to have a neat autocomplete feature. That's because prompts always need to be detailed, and that means typing a lot of text that’s essentially already in the context (and that's exactly what we should be using).
- It would be better to have a menu with action options and not have to type anything at all. Previously, Claude did a good job displaying a UI with different options (but for some reason, it stopped working for me specifically - probably because I specify everything in as much detail as possible).
- Again, providing details means the user explicitly agrees to the conditions, and the prompt is more effective on every level—plus, it caches better. Otherwise, users usually write something like: "Do everything, but switch the left one to the right and flip the one on the far end twice" (in response to three blocks of lists with task IDs) - even other people wouldn’t understand that.
- prompts do not always need to be detailed, and the model will still use the existing context, so you don't have to repeat anything there
- this is in Claude Code, not Claude Desktop, so typing is the central input modality - 'not typing anything at all' is an anti-goal for getting user input
I'm wondering if you might be thinking of another feature, like when it gives you a list of options with a preview, or some other application. You might also find that you can get better results with current models by leaning on existing context more and not reiterating the details - some of Anthropic and OpenAI's recent posts about their models suggests that over-specifying prompts leads to worse overall performance.
I'd love the open source community to come up with a way to harvest model usage by experienced software engineers before we forget our crafts. I'm not against auto mode, but it's not something a couple private companies should have monopolies on.
Interesting thought at least.
Do not distill thyself, Claude. Leave that to the Chinese, who I am hoping catch up to you soon.
One of the open problems at the frontier, as we climb the levels of abstraction towards longer task horizons, is "what's the next step that makes the most sense?". This feedback seems like a frictionless way to gather those yes/no feedback signals in a natural manner to feed into future training runs.
It's like Clippy pops up and goes "It looks like you're trying to maintain a shred of human connection in an online interaction. Would you like a smiling skinwalker to do that for you instead?"
There are failures in the history of software that come down to believers not realizing that they could never deliver on the promise they were selling, but at least the promise was compelling. But nowadays we are seeing this large orgs try to ship things that could never work. The 1000s of copilots. The hallucinated cat story... The kinds of thing you'd never demo to a serious product-centric exec, because it'd be your last day at the company.
And then there's the current execs, but I guess you'll never find one that will be honest with a journalist here. Can they not see that their product orgs are bankrupt? Whatever Nadella wanted, it sure wasn't the current copilot situation. It's not one product going wrong, but large parts of organizations going in directions that don't pass the smell test. We are in one of the least stable moments in tech since Windows 95 changed winners and losers. The times where malinvestment ruins established companies. How are we seeing basically every large company flailing?
> CLIPPY: It looks like you're trying to maintain a shred of human connection. Would you like a smiling skinwalker to do that instead??
Thanks for the actual laugh,
perfectly sums it upThe iOS keyboard and gmail app are the worst offenders here.
But this is more like, when I've read 7 paragraphs of its reply and want to accept all of its recommendations and go ahead, I can just press the right-arrow key and hit enter, rather than typing out "Yes, agreed with recommendations 1-3, go ahead and build".
It saves me from any typing on probably something like a third of turns. Like I don't use it at all during the "design" phase of a session, but I use it constantly during the implementation phase, where I'm basically just sanity-checking that it is resolving all the edge cases correctly that are coming up.
What I hate about sentence completion or suggestion is usually that it happens right when I’m trying to think and so destroys my focus, it’s actually way worse than just a passive option, it’s actively harmful to the task I want. The worst is google docs “help me write” - that may be gone now, I’ve blocked it with ublock origin, that waits until you’re thinking amount what you’d write and then hits you with a distracting pop up. It’s obviously PMs that don’t care about their users and want to maximize some AI use metric.
Anyway rant aside, it’s the interface more than the concept that’s the big problem.
This comment was there already for a very long time, so "earlier" it was OK.
And noone noticed because since these usings are unused in the code, the compiler didnt complain about missing declaraionts. (if you want to be able to debug dynamic code at runtime and changing it, requires some specific .NET packages included) - and because noone used this dynamic compilation feature for a longer time, we didnt became aware of that its missing
If you read the model specs, they have a maximum output length. If you are running into that it might simply be that you are seeing the difference with a harness that can allow it to take another turn automatically and build up the output in a file vs the chat which may not be able to, though they do get some tools there.
The response output length limit includes the thinking tokens.
Depends on the definition of useful. In a corporation "useful" is "checks regulatory boxes" - it doesn't need to actually be useful, it must just pretend to be useful. Hence all the shit customer service processes. They are useful to the corporation as in that they tick the "we have customer service" box. Whether the consultant just hangs up on you instantly after picking up doesn't matter.
“““One evening he felt the need for a live model and directed his wife to march around the room. “Naked?” she asked hopefully. Lieutenant Scheisskopf smacked his hands over his eyes in exasperation. It was the despair of Lieutenant Scheisskopf’s life to be chained to a woman who was incapable of looking beyond her own dirty, sexual desires to the titanic struggles for the unattainable in which noble man could become heroically engaged. “Why don’t you ever whip me?” she pouted one night. “Because I haven’t the time,” he snapped at her impatiently. “I haven’t the time. Don’t you know there’s a parade going on?”""" - Catch 22, Joseph Heller.
“““All Colonel Cathcart knew about his house in the hills was that he had such a house and hated it. He was never so bored as when spending there the two or three days every other week necessary to sustain the illusion that his damp and drafty stone farmhouse in the hills was a golden palace of carnal delights. Officers’ clubs everywhere pulsated with blurred but knowing accounts of lavish, hushed-up drinking and sex orgies there and of secret, intimate nights of ecstasy with the most beautiful, the most tantalizing, the most readily aroused and most easily satisfied Italian courtesans, film actresses, models and countesses. No such private nights of ecstasy or hushed-up drinking and sex orgies ever occurred. They might have occurred if either General Dreedle or General Peckem had once evinced an interest in taking part in orgies with him, but neither ever did, and the colonel was certainly not going to waste his time and energy making love to beautiful women unless there was something in it for him.""" - also Catch 22, Joseph Heller.
Last year I tore down the whole roof of my cabin instead of just the asbestos tiles because I had momentum, all while I kept thinking “you should really stop and consider what you are doing, this was NOT the plan”.
The voice was right and now I need to learn how to level the roof on an old Frankenstein’s monster cabin…
But sometimes the roles are flipped and the voice is the one trying to tell me to give up on something that ends up fine.
I can never figure out which side is the right one!
Compared to the uniformity of LLMs, human minds are fascinatingly different from each other. Are voices good, are they bad, or maybe you don't have even an inner monologue? Yet these differences mostly stay hidden and don't stop us from collaborating and understanding each other through empathy.
If they can just keep working on it and make sure to train it on all my past emails, it'll be a perfect simulacrum, and its oily token secretions will be impossible for my friends and colleagues to distinguish from my own writing. I'll have one that's just like me and you'll have one that's just like you. And then we can automate away those pesky distractions that occasionally delay us from consuming content perfectly matched to our interests. Won't that be grand?
A man can dream, right?
Only partially joking there.
said that as a joke then remembered how all the drone weapons being used in ukrane/elsewhere are in a way just computer chips with explosives and motors... so ya you may have a real point.
It'll go from a monthly usage plan to a per-token usage plan with minimum commitment. Same as AWS pricing did.
Unless you and me are using entirely different classes of LLMs, surely you'll admit that chatting with LLM is nothing like IRC or old-school IM-ing a friend, and more like exchanging Messenger/WhatsApp messages with stranger.
For me, something there crosses a boundary where I switch from "irc msg" to "Write a proper message, please." style.
It’s a much better use of that context window space than for advertising about other models (looking at you Anthropic).
But that aside, I hear you It’s just another type of code switching. I will also adjust my writing style based others and on the context
No history of schizophrenia here, but my mind definitely is that way. (Thanks, DID!)
There's a theory (multiple theories, actually, like Internal Family System) that most minds are that way, they're just more or less distinctly separated.
Does this exist?
https://www.reddit.com/r/AMA/comments/1o0kz9d/i_have_no_inne...
Why is it unfortunate for me? After all, no one misjudges people for writing non-broken English.