Cookies

We use analytics to see how the site is used so we can improve it.

Skip to content
Renada

Why Claude costs double GPT5 but wins every coding benchmark

A plain English comparison of GPT5, Claude Sonnet 4.5 and Gemini 2.5 Pro for MSP owners deciding where to start with AI.

30 October 2025 15 min watch Connor Fagan

The short version

Connor from Renada opens a new series on AI for MSPs by comparing three flagship models: GPT5, Claude Sonnet 4.5 and Gemini 2.5 Pro. He breaks down context windows, pricing per million tokens and coding benchmarks so you can pick the right model for the job rather than assuming one model does everything well.

What you'll take away

  • The contractor analogy

    Different AI models exist for the same reason you would not hire a joiner to paint a delicate mural. Pick the specialist for the job, not a generalist by default.

  • GPT5 base plan gives you 8,000 tokens

    That is roughly 10 pages of context before the free web version starts to lose track of what you told it.

  • Claude Sonnet 4.5 scores 77 to 82 percent on SWE benchmarks

    It is the clear leader for coding work, well ahead of GPT5 at 75 percent and Gemini at 64 percent.

  • Claude is the most expensive of the three

    3 dollars per million input tokens and 15 dollars per million output, roughly double what GPT5 charges.

  • Gemini 2.5 Pro remembers the most

    A 1 million token context window, with a newer model pushing to 2 million, equivalent to about one and a half thousand pages of text.

  • Claude code can run autonomously for hours

    Anthropic claims up to around 30 hours of autonomous work using proper development tooling like Claude code, not the standard chat interface. Connor uses it for elegant insights because it writes strong SQL.

Key insights from the episode

  1. Context window size directly drives hallucination rate, a bigger window means the model forgets less of what you already told it.

  2. GPT5's free web tier gives only 8,000 tokens of context, upgrading to the pro plan or the API unlocks far more.

  3. Claude Sonnet 4.5's 200,000 token window can be extended to 1 million via the API, but this costs extra.

  4. For coding tasks pick Claude Sonnet 4.5, it leads all benchmarks Connor references and is the model behind elegant insights.

  5. For large document processing or anything inside Google Workspace, Gemini 2.5 Pro's context window and native integration make it the safer default.

  6. GPT5 and Gemini both charge 125 dollars per million input tokens and 10 dollars per million output, Claude charges 3 dollars input and 15 dollars output per million.

  7. Claude cannot generate images at all, while GPT5 and Gemini both handle image generation well.

  8. Privacy policies differ and change quickly, check the current terms for each vendor rather than assuming API traffic is never used for training.

Questions people actually ask

Which AI model is best for coding, GPT5, Claude or Gemini?

Claude Sonnet 4.5 is the strongest coding model, scoring 77 to 82 percent on real world software engineering benchmarks compared with 75 percent for GPT5 and 64 percent for Gemini 2.5 Pro. It is also the model Connor uses for writing SQL in elegant insights.

What is a context window in AI models?

A context window is effectively the model's short term memory, the amount of text it can hold in mind during a conversation before it starts to lose track or hallucinate. GPT5's free web version gives around 8,000 tokens, Claude Sonnet 4.5 gives 200,000, and Gemini 2.5 Pro gives 1 million, with a newer model reaching 2 million.

Is Claude Sonnet 4.5 more expensive than GPT5?

Yes, Claude Sonnet 4.5 costs 3 dollars per million input tokens and 15 dollars per million output tokens, roughly double GPT5's 125 dollars figure quoted per million input and 10 dollars per million output in the video. It is the priciest of the three models Connor compares.

Can Claude generate images like GPT5 or Gemini?

No, Claude cannot generate images at all. GPT5 and Gemini 2.5 Pro both handle image generation well, so pick one of those two if that is your use case.

Which AI model integrates best with Google Workspace?

Gemini 2.5 Pro, because Google owns both the model and the Workspace apps, so it interacts with Docs, Sheets and Drive more reliably. Other vendors can connect too, but Gemini is the safer assumed bet for that ecosystem.

Why does GPT5 hallucinate more than Claude or Gemini?

GPT5's smaller context window, especially the 8,000 token free tier, means it forgets earlier parts of a conversation sooner, which increases hallucination. Claude's 200,000 token window and Gemini's 1 million token window hold far more context, which noticeably reduces this problem.

How long can Claude Sonnet 4.5 work autonomously?

Anthropic claims Claude Sonnet 4.5 can work autonomously for up to around 30 hours when used with proper development tooling like Claude code, not the standard chat interface. Connor describes running multiple agents at once with Claude code as genuinely impressive, though it comes with tradeoffs on cost.

Full transcript

2,904 words

Read full transcript

Connor: Every event you go to, every post on LinkedIn, they're all telling you that you must use AI within your business. But do you? And if you do, where do you start? Or how do you start? Well, fear not. Renada has intentionally not put any content out for the last two years. I've been avoiding the comments like this and this and these for years because I've needed time to understand AI. I needed time to understand the impact on your business. I mean, it's time to see where the dust settles and to really qualify what I'm about to say. So, this has been the most stressful video to create to date. And this is me just trying to explain AI because it's complicated and I need to ensure the things that I'm saying are factually correct.

The irony is I recorded a video two weeks ago and all the models change. So, we're back here doing it again because I wanted it to be right on the time of recording. But anyway, we're going to do a small series all about AI. We're going to be chatting about technical jargon such as tokens and prompt caching and hallucinations and RAG and vector embeddings and all the other things and I'm going to break them down into English for you or at least I'm going to try. The vector embeddings video, you're going to have to just bear with me. So, my name's Connor Solutions and I'm going to hopefully help you understand AI and remove all that overwhelming fear that could be associated with it. So, let's get stuck into it. Section one is all about models and why or why not you might pick those.

So, welcome to section one. Today we're going to be talking about different models. What is a deciding factor for a model and why do we even have different models and which one should you use? Typically, it's a very difficult place to start. So, I'm going to start with a good old analogy. The reason we have different models is the same reason you wouldn't hire the same contractor to renovate your entire house. You would hire a plumber to do the plumbing work. You would hire electrician to fit your lights and your sockets. And you would hire a joiner to do all the joinery work like fitting your doors. Right now, it's obviously fair to say that a joiner could paint your walls for you. However, a painter or decorator would probably do a better job. However, a better job always comes at a cost. And do we need to spend the additional cost for the output?

I'm sure anyone could paint a wall white, but if you're wanting some delicate artwork doing, then you might really want to pick a very detailed painter. So, we're going to chat about three different models today. We're going to talk about GPT5 by OpenAI. We're going to talk about Claude Sonnet 4.5 by Anthropic. And we're going to talk about Gemini 2.5 Pro by Google. So just a quick thing to understand, different vendors have loads of different models within them. So I'm really going to focus on today OpenAI, Anthropic, and Google as the top three vendors. There's loads of vendors out there now, but I'm just going to stick with those three. And then we're also going to pick what I think is their flagship or best model from underneath it, which is GPT 5.0 currently, Sonnet 4.5, and Gemini 2.5. Now, there are loads of other different models within each of those vendors, but we're going to ignore those for the time being because this video would never end.

So, let's start with ChatGPT5. OpenAI was the first accelerant, I think, for AI in the modern world as we know it. ChatGPT, the model way, way way back when, was the first one to come to market for consumers really. It's what most of us know. I would argue it's probably the most popular as well. But ChatGPT5 specifically was released in August of 2025 this year and it's their latest flagship model as of the time of recording. And it's their most versatile model. It handles everything pretty well from writing to coding to maths to healthcare questions. It's a very good all-rounder.

Secondly, ChatGPT5 has significantly reduced hallucinations. We're talking about 45% fewer errors. We will also, by the way, put loads of references to some of the stuff I'm speaking below so you can fact check what we're saying. I think it's important on this subject. It's also, I suppose, coming out. It's now got a bigger context window. We'll talk about some of the comparisons shortly, but ChatGPT for me always hallucinated really badly, and that's because the context window was always really small. It would forget what you just told it as soon as you just told it or at least that's what it felt like. So that's what we're going to talk about in a minute is ChatGPT5.

In the free version, by the way, I think you get 8,000 tokens of context. That's about 10 pages before it starts to hallucinate or if you're doing a lot of depth, then it, you know, soon ends. It's also not the absolute best at a singular thing. It's good at everything, but their specialist area isn't really defined. Personally, I think it's pretty good at image generation, but we'll stay there. And the pricing, just for the screenshot here, it's around £1.25 per million input and around £10 per million output. So, it's very mid-range in the pricing bracket. So, that is ChatGPT5.

Now, let's talk about Claude Sonnet 4.5. This was released by Anthropic very recently. So, Claude Sonnet 4.5 was released just a few weeks ago at the time of recording, so in September 2025 this came out. And here's the thing, if you're doing any kind of coding, this is the model you want in my opinion. People will argue ChatGPT5 is just as good. I think they're lying. We'll give some facts on this soon, by the way. But for me, it's the flagship right now. It is the model for coding.

It's also the reason we use it for elegant insights because it's really good at writing SQL and I still think to this day it's the best model. But first it is right now by most benchmarks the best coding model in the world. It scores 77 to 82% on real-world software engineering benchmarks. That's significantly better than all of the competition. We'll put in some of the SWE benchmark links below so you can go validate this yourself. Second, and it's quite remarkable, when using proper development cycles like Claude Code it can work autonomously for hours and I think Anthropic claim it's like 30 hours by the way. It can work up to but let's not get into semantics. But to be clear this isn't regular chat interface. This requires specialised tools for developers building complex applications. But it is game-changing. We use Claude Code in development and when you get multiple agents going at the same time it is quite unbelievable. But there are some tradeoffs.

Claude Sonnet 4.5 to me is quite an expensive model. It is £3 per million input and it is £15 per million output. That's about double-ish what ChatGPT5 charges. The context window though is 200,000 tokens. This means its hallucination is a lot lower. You can even, I think, increase now over the API to a million input tokens if you really want to. And that's not available for similar ChatGPT inputs. I think ChatGPT we said earlier was 8,000 for the chatbot online whereas Claude is 200,000 input tokens. So, massive variance there. You will see the hallucination difference in those models as you are working through it. To be clear, if you're coding, Claude Sonnet 4.5 is definitely your pick.

And then we have Google's Gemini 2.5 Pro. Gemini 2.5 Pro is Google's flagship thinking model. First, it has the largest context window by far. That's 1 million input tokens, by the way. And I think either coming soon or is out now, there's the new model, which has 2 million input tokens, which is just insane. And that's like one and a half thousand pages of text, by the way, that it can remember. So, context window is memory. We'll touch more on this later. It integrates really, really well with the Google Workspace. So if you live in Google Docs or Sheets or Drive, this model can interact with them and access them really, really well. Now other providers like Anthropic can do this as well and I think ChatGPT can, but if you think that Gemini owned by Google can integrate with Google, it's typically a safer bet or assumed at least that it should be better.

For coding though, it scores around 63.8% on the SWE benchmark. So, it's decent, but definitely not a flagship or the best. Also, when you're using thinking mode for Gemini, it's around 30 seconds for your first response, which is quite slow if you're thinking about it. So, if you're processing massive documents and need a cheap option for it, then think about the Google ecosystem for that. You might get really, really good results.

So, that's kind of three models as a summary. Now, let's get some side-by-side comparison. And Dylan's going to do some magic work now on screen between here and a table or put me down. I don't know. We'll sort it out. Let's do a side-by-side comparison. What is the difference between the three different vendors and specifically the three different models that we're referencing today? Now, let's put these side by side and let's talk about context window. So, GPT5 context window is 8,000 to 400,000 tokens depending on how you access it.

Now, again, this is really, really, really important to understand. If you're just using ChatGPT via a web browser, I think depending on the plan you're on will dictate how many tokens you get. So, I think the base plan is 8,000 tokens, which is tiny in my opinion. But if you're on their Pro £200 plan, then you can get more and then if you use the API, you can enable it to do even more. So, it's complicated is what the point I'm trying to make there. If you're using Claude Sonnet 4.5, well, you're getting 200,000 tokens unless you enable via the API when you can then get a million tokens. But again, you've got to tailor it. Just to be clear, by the way, with Claude, if you do enable the million token option or add-on, it does cost more money by the way. So, go and check that out yourself.

Gemini though, 1 million input tokens or 1 million context, should I say, tokens, which is massive. It blows all the competition out the window. So what this means is the model GPT5 will remember less than Gemini. That's the takeaway.

Then let's talk about pricing. So, GPT5 is £1.25 for a million input. Gemini is also £1.25 for input, but Claude Sonnet 4.5 is £3 for 1 million input, which is the most expensive. Then let's talk about output tokens. ChatGPT and Gemini are both £10 for a million output tokens where Claude is £15 for a million output tokens. So, the most expensive for raw input output tokens. There will be a full video by the way explaining tokens and what they mean. But in essence, when you're sending messages to AI versus getting stuff back, that is input and output.

Let's talk about use cases quickly. So, GPT5, it's an all-rounder, can be fairly good for healthcare apparently, just a general purpose model. Claude is amazing for coding and software development and autonomy. You can really get agents with Claude to be autonomous. Gemini, what's that good at? Well, it's good at really large document processing because of that massive context window. It is good for research and in my opinion the best for integrating into Google products.

Then let's talk about coding performance. If you're a coder, we are technical nerds or typically that's a demographic. So what's a good use case for that? Well, if we look at the SWE benchmarks as a starter, or we'll actually put benchmarks down below actually as a whole thing, but Claude wins at 77 to 82% in benchmarking for coding. GPT5 is now up there on 75%. And Gemini is third at 64%. So Claude for me is the winner on coding.

Privacy. Now, this is getting complicated because Claude has just changed or Anthropic has just changed their privacy model. I'm just going to put complicated for them all if I'm being honest. You really need to look at the privacy policies of these models. API routes, really need to be looked at. I don't think any model learns via the API or you can opt into it. Now, if you're using the consumer front end, so the web page, I think they're now all locked out. But again, please do look at the privacy policies for these models. I'm just going to put complicated in the table below.

So, that's a quick highlight really of kind of where they bench against each other. It's complicated is the long and short of it. What I would say is play with the models, get a feeling for what they do well for you. For me, for instance, ChatGPT is great at image generation, but so is Gemini now. Claude can't do it, right? It just can't do it. But Claude now can integrate well and do word document generation and can read Excel files. So, there's loads of different variants with the models. I'm not going to go through them all today because I'll put you all to sleep. But please do test different models by different vendors and see what's worked best for you in your journey of doing this.

So, what's the conclusion then? Three models to remember. ChatGPT 5.0 is the flagship right now from OpenAI. Claude Sonnet 4.5 is a flagship by Anthropic and Gemini 2.5 Pro is the flagship from Google. Go and research what models you need. Also throw the claims into the models themselves. Get different models to validate different models and it is generally pretty good at that. So, I use Reddit a lot. I look at Discord channels. I play with them all.

This video has been a lot longer than I wanted because it is such a complicated thing to unravel in one's mind. There will be a bunch of resources below. So, benchmarking, genuine claims from different research articles that I want you to go and read and look into. It's complicated. I probably made it way more complicated because that is the nature of this thing. But I hope this starts. In the next video, we're going to start talking about some jargon. So, what does context tokens or context window actually mean? What do tokens mean? What does prompt caching mean? So, please do subscribe and like this video if you like what I've been rambling on about. I've been Connor Fagan. I hope you found this useful. Have a beautiful day and I will see you all soon for the next video. Take care. Goodbye.

Great experience with the Renada team. Not too often you find people who are not only competent but enthusiastic and genuinely invested in your success.
Netaryx Google Logo

Our Core Services

Offering support to enable sustainable success for your organisation.

Consultation Harness the transformative potential of an agnostic advice tailored to your unique business needs. From PSA implementation to ongoing support, our exceptional consultation services pave the way for extraordinary success. Find out more
Virtual Admin Let us handle the technical heavy lifting. Our expert team builds solutions, creates powerful reports and dashboards, and develops automated integrations - giving you more time to focus on what matters most: your clients. Find out more
Product Onboarding We understand that the first steps in adopting a new product can be daunting, we are here to guide you through every stage of the process with precision and clarity. From initial setup to advanced features, maximise the value of your product from day one. Find out more
Virtual Chief Technology Officer (vCTO) Benefit from a remote and adaptable technology expert to seamlessly combine strategic guidance and effective leadership to propel your business to new heights and empower your organisation’s technology ability. Find out more
Where to next? Get the cutting-edge tools to support your MSP business. Contact us today to receive a bespoke quote tailored to your specific needs.