I don't get it. I like DeepSeek, because I can turn on Search button. Turning on Deepthink R1 makes the results as bad as Perplexity. The results make me feel like they used parallel construction, and that the straightforward replies would have actually had some value.
Claude Sonnet 3."6" may be limited in rare situations, but its personality really makes the responses outperform everything else when you're trying to take a deep dive into a subject where you previously knew nothing.
I think that the "thinking" part is a fiction, but it would be pretty cool if it gave you the thought process, and you could edit it. Often with these reasoning models like DeepSeek R1, the overview of the research strategy is nuts for the problem domain.
O1 doesn’t seem to need any particularly specific prompts. It seems to work just fine on just about anything I give it. It’s still not fantastic, but often times it comes up with things I either would have had to spend a lot of time to get right or just plainly things I didn’t know about myself.
I don’t ask LLMs about anything going on in my personal or business life. It’s purely a technical means to an end for me. So that’s where the disconnect is maybe.
For what I’m doing OpenAI’s models consistently rank last. I’m even using Flash 2 over 4o mini.
I'm curious what you are asking it to do and whether you think the thoughts it expresses along the seemed likely to lead it in a useful direction before it resorted to a summary. Also perhaps it doesn't realize you don't want a summary?
Interesting thinking. Curious––what would you want to "edit" in the thought process if you had access to it? or would you just want/expect transparency and a feedback loop?
I personally would like to "fix" the thinking when it comes to asking these models for help on more complex and subjective problems. Things like design solutions. Since a lot of these types of solutions are belief based rather than fact based, it's important to be able to fine-tune those beliefs in the "middle" of the reasoning step and re-run or generate new output.
Most people do this now through engineering longwinded and instruction-heavy prompts, but again that type of thing supposes that you know the output you want before you ask for it. It's not very freeform.
If you run one of the distill versions in something like LM Studio it’s very easy to edit. But the replies from those models isn’t half as good as the full R1, but still remarkably better then anything I’ve run locally before.
I ran the llama distill on my laptop and I edited both the thoughts and the reply. I used the fairly common approach of giving it a task, repeating the task 3 times with different input and adjusting the thoughts and reply for each repetition. So then I had a starting point with dialog going back and forth where the LLM had completed the task correctly 3 times. When I gave it a fourth task it did much better than if I had not primed it with three examples first.
Claude Sonnet 3."6" may be limited in rare situations, but its personality really makes the responses outperform everything else when you're trying to take a deep dive into a subject where you previously knew nothing.
I think that the "thinking" part is a fiction, but it would be pretty cool if it gave you the thought process, and you could edit it. Often with these reasoning models like DeepSeek R1, the overview of the research strategy is nuts for the problem domain.