Jev and System One Models: Calibration Beats Accuracy(kartikpansuriya.com) |
Jev and System One Models: Calibration Beats Accuracy(kartikpansuriya.com) |
I'm not sure that's a good way to use Jev for this use case. If I had this problem, I would ask Jev to answer several different yes/no questions about each message, and then use the probabilities as inputs into a logistic regression model that predicts urgency.
If you ask the model specific yes/no questions which can be answered reasonably objectively from the input, I think the answers are going to be more stable over successive generations of models.
e.g. if you ask 'Is the customer angry?' I'd expect that answer to have high agreement between models and between models and humans. But directly answering the 'is it urgent' question is much harder. (Although I suppose you can try to put the rules in the prompt.)
This is the part that bothers me the most. How is it 0 hallucination, if the correct answer is not even the part of the options. There is no way to mark absentia or a way to know I absolutely cannot choose any of the options.
I am wondering, if anyones tried dead simple combinations of embedding with logistic regression to solve classifications problems?
It would be absolutely hilarious if you did this, asked it which model it was, and it gave a recognizable answer that wasn’t Jev.
Edit: Dude, you didn’t even use the thing you’re talking about? It’s on OpenRouter. Do better!
Edit 2: OP is a ~60 day old account, only other (positive) commenter is a ~48 day old account. Sus.