Replies: 3 comments 4 replies
|
I think 5.5 was very experimental , considering they only released a single version I also really like the way it format words, almost to the same level of claude models (I will repeat myself 5.5 and 5.4 is better than claude-sonnet5) but for the way it speaks and talk, 5.6 is on another level im not a big fan of but its not like it used a command that wouldnt work actually , Ive put put 5.6 to the test, and it actually used a library and tool --- while 5.5 didnt used a tool and it didnt tried to reinvent the wheel and analyse the a literal .c10 binary so very good! I honestly was so disapointed at 5.2 and 5.3 (but I like the way it tries to talk with you regarding task), 5.4 was truely something else , but I understand why, because the older models didnt really compete with the other models despite that I saw the pontential because for using windows powershell at the time only openai done it correctly so yeah very good update, considering things are now faster ---- and I think it choses things that are good and things that the users would be happy (you dont need to write a C++ compiler for everything!) |
|
Honestly I think the new naming changes is great, I think the users should refain from use the bigger models all the time , when they can be just be happy with the smaller models. and to be honest, I am sure in a lab testing openai can achieve even 80% where Sol 5.6 eached on evals 60%... (not-or 0-shot, or very expensive even in lab setting) and also I guess again , I hope people try to expirement with the smaller models first and try to actually adjust the smaller models like tera or luna, and not go full power (because unlike other companies, if you choose the bigger models, they do not degrade your service to the same degree more than 30%-10% , which is not 65%) I personally wasnt a fan of Claude Code, because it took 5 seconds to understand at december 2025 that their documentation is cheotic (more than my writing style yes) and so much stuff in their documentaiton gets aboanded in like 3 days as per Klaus Claude, recently that launch a feature called ULTRATHINK, and it would literally will launch like 105subagents or whatever, and it would burn into your 5-houry credit in 3 prompts on a small [emtypical-vibe-coded small code] , how dare they release this to users? also burning all of this compute when they clearly cant serve as things going now as per subscribtion side , I also think if I paid in 2024 for few months in a faw for a claude subscribtion , I do hope to get basic prompt ablity ,, because once I am not a paying customer and I became a free users--- I I should become like free_paid_in_the_past and get a very basic preview, but instead I used other third party services - which lets be honest poision their data -- and anthropic doesnt enjoy the same kind of luxury of most people using MicrosoftCoppilot or Codex, and a lot of the things regarding how messages are handled happen quickly openai even changed the their standards regarding text/agent-chatting -- anthropic has a lot of data from a lot of third party apps --- which I assume they can literally use llm based apporch to even clean the dataset (find start, and end goal --- or if the user achieved what he wanted) as per raw chat, I actually thing it can get creative, but I think it get posion with a massive context window you already load before a you start a new chat (system card? is that that?) I actually think that adding a lot of guildlines and safety or whatever are good for the models, because it adds a lot of complexity on how the models themselves are going to desgined and I think its part of the reason it got driven so fast in the last two years -- because it does degrade preformance and it forces team to expirement more and more (but thats just a theory of mine because I believe you should break your head with a lot of limitation) lets real there is more than just matrix multipications , every itiration, would provide much more rather than labeling it as a data problem and pass it to the cpu, |
|
I tell you what you wont be disappointed with the smaller models - I dont think you will tell the difference, despite what you choose is what you get --- (that cant be said on other companies) |

Uh oh!
There was an error while loading. Please reload this page.
Just wanted to say, i've only been using it for a couple hours and I was worried the Codex mode was still going to bulldoze instructions and read paths from the efficiency methods and parallel reads, but it seems to follow instructions well. Been running some rigorous tests across all the models (only finishing up Sol right now) but so far so good. And High performed very well. Started off to a rocky start with xhigh burning through 41% of my usage window when i asked it the first question in the chat about the difference of token window usage from 5.5 Codex to 5.6 ChatGPTWork. Did not expect that haha. I'll keep testing but wanted to say thanks to your team.
All reactions