Cookies

We use analytics to see how the site is used so we can improve it.

Skip to content
Renada

Why your AI bill hides in tokens, caching and context windows

A plain-English walkthrough for MSPs trying to understand what tokens, RAG, embeddings and hallucinations actually mean before they wire AI into HaloPSA

28 January 2026 19 min watch Connor Fagan

The short version

This episode of Renada Rundown decodes the AI jargon MSPs keep bumping into: tokens, context windows, prompt caching, RAG, vector embeddings, structured outputs and hallucinations. It uses real examples from Renada's own product, Elegant Insights, and from HaloPSA auto triage to show why each term actually matters for cost and reliability, not just theory.

What you'll take away

  • Tokens are the currency of AI

    You pay for input tokens and output tokens separately, and output is usually far more expensive per million than input.

  • Claude Opus versus Claude Sonnet, a real price gap

    Opus runs around $15 per million input tokens and $75 per million output tokens, while Sonnet sits at roughly $3 and $15.

  • Prompt caching turned a break-even product profitable

    Elegant Insights was sending a 60 to 80 thousand token database file on every message, costing around 30 cents a time, until Claude's one hour prompt caching cut that to a fraction of the price.

  • Context window size decides how much HaloPSA data an AI can hold at once

    ChatGPT's typical 128k token window is roughly 200 pages of text, against Gemini's 1 million token window at around 1500 pages.

  • RAG is just AI searching your documents

    Retrieval augmented generation means the AI checks your knowledge base for a similar ticket instead of guessing or scraping the open web, where AI generated slop is a growing risk.

  • Structured outputs stop auto triage from guessing the format

    Asking AI for a category, priority, impact and urgency only works reliably in HaloPSA if you force a JSON response, otherwise the fields come back in a different order or go missing.

Key insights from the episode

  1. Input tokens cost far less than output tokens on most models, so long AI-generated replies cost more than long prompts sent in.

  2. Roughly 100 tokens equals about 75 words on average, but short or unusual words can each count as a single token.

  3. Use a token counting plug-in or API endpoint against your chosen model to work out real input and output token counts before estimating cost.

  4. A 1 million token context window costs noticeably more than a smaller one, so match window size to how much data you actually need retained.

  5. Force structured outputs, ideally JSON, whenever you feed an AI response back into a system field, such as HaloPSA ticket category or priority.

  6. Skip structured outputs for free text tasks like rewriting a ticket reply, where some creativity in the HTML response is fine.

  7. Keep a human validating AI-set ticket fields, since models can hallucinate, drift, or degrade in performance on a given day.

  8. Temperature controls creativity from zero, strict and repeatable, up to one, freewheeling, but even zero temperature never guarantees an identical answer twice.

Questions people actually ask

What are tokens in AI and why do they matter for cost?

Tokens are the units AI models charge for, split into input tokens for what you send and output tokens for what the AI replies with. Output tokens are typically priced far higher than input tokens, so verbose AI responses cost more than long prompts.

What is a context window in AI models?

A context window is the amount of data a model can hold and process in one go, measured in tokens. ChatGPT typically offers around 128k tokens, roughly 200 pages of text, while Gemini can handle up to 1 million tokens, around 1500 pages, letting it analyse entire log files or multiple tickets at once.

What is RAG in AI and how does it work with a helpdesk?

RAG, retrieval augmented generation, is AI searching your own documents rather than making an answer up. A ticket like VPN keeps disconnecting gets matched against your knowledge base so the AI answers from your documented solution instead of guessing or pulling from the open web.

What is prompt caching and why does it reduce AI costs?

Prompt caching lets a model store a large chunk of repeated context, such as a database file sent with every message, so it does not need reprocessing at full price each time. Claude's move to one hour prompt caching cut Renada's Elegant Insights running cost enough to turn a break-even product profitable.

What are vector embeddings and how do they relate to RAG?

Vector embeddings turn your data into a string of numbers based on context, not just the literal words, so similar meanings sit close together numerically. When a new ticket comes in, it gets embedded too and matched against existing embeddings, essentially a game of snap scored by percentage similarity.

Why do structured outputs matter when using AI for ticket triage?

Structured outputs force the AI to return data, such as category, priority, impact and urgency, in a fixed format like JSON. Without that structure the AI might reorder fields or miss one out entirely, which breaks any automation trying to update a ticket from the response.

What causes AI hallucinations and how do you reduce the risk?

Hallucinations happen when an AI drifts, runs low on context, or simply makes something up, and they can vary between models or even between runs of the same model. Keep a human validating AI output before it reaches clients, and treat temperature as a dial between strict and creative rather than a fix on its own.

Full transcript

3,692 words

Read full transcript

Welcome back to this AI miniseries this month. In the last video, we had a quick intro and spoke about different models and why you may or may not want to use them. In this video, we're going to get into some of the jargon. I'm going to decode hopefully some of those words you've been seeing so you can kind of understand what the hell you're doing with it. So let's get stuck into it.

Let's talk about the most important thing you're going to want to understand and that's tokens. What the hell are tokens? Well, tokens are the currency of AI. Whenever you use an AI model, you get charged for input. I know we'll touch on that in a minute. And you also get charged for output tokens. So when you send a message to an AI, you get charged for that message. It's a point of contention. Don't get me started. Then it processes it. And then when the AI responds, you get charged for that as well.

Now, the good thing is most models the input is way way cheaper than the output. I'm assuming that's because of the processing power. It's the way they price it. If you're sending a few words, it's really cheap. If you're sending essays and essays and essays for AI code, it can get quite expensive.

Depending on the model you use, and I say model very particular, depending on the model you use will dictate the cost. So let's take Claude for instance. Claude, or Anthropic, should I say. Anthropic has Claude Sonnet 3.5, has Claude Opus, and Claude Haiku. I think they've renamed them all to Claude now. Anyway, let's say three models. Opus is their most creative model apparently and this is vastly expensive. Something like £15 per million input and £75 per million output, I think. Claude Haiku is their really really cheap model. I don't have prices to hand, look at them. And then Claude 3.5 Sonnet is their mid-tier model. That is £3 for a million input tokens and £15 for a million output tokens. If you compare that to something like ChatGPT, that is way cheaper. However, you get what you pay for sometimes.

When I say sometimes, different models perform, as the last video, better at doing different things. So if you're asking Claude Opus to write you a LinkedIn post versus ChatGPT 4, you may find the outputs are very similar, but the cost difference will be huge. So how does the costing actually work? Well, costing is based obviously on the tokens. And the way we determine what an input token is is a little bit gray. It's not quite one token per word. And it's also not based on the number of characters in a word. It's kind of both.

Let me give you an example. A hundred tokens is on average around 75 words. However, if you're just doing like, I don't know, "a a a" which is four letters, that may very well just be a single token. There are loads of token calculators online for you to go and look at and work out. I believe each model is slightly different. So when we're building in Claude, they have a token counting plugin you could use, an API endpoint you could use to track your tokens. So you can actually work out how many tokens you're using input and output and then try and work out the cost from that.

With tokens you've also got something called context window. And the bigger the context window, the more data the AI can store and process for you. So ChatGPT is typically about 128k tokens for an input, where Gemini is 1 million tokens for an input. This means you can send a million tokens worth of data to the AI and it can process it all quite nicely. Claude actually brought out recently that there's a new model or you can leverage a 1 million context window, but that costs a lot more money. So depending on how much data you're sending to AI or you're having a conversation, how much data you want it to retain, your context window really matters.

And context window, for instance, 128,000 tokens is around 200 pages, or 1 million tokens for Gemini is around 1500 pages worth of text, just as some context for you there. So bigger window, analyse entire log files or multiple tickets all at once as opposed to a small window where you may have to do loads of separate calls and then ask it to analyse the output, if that makes sense.

So tokens is the currency of AI. When we're looking at tokens and cost analysis and things like that, we want to be mindful of caching. So certain models allow you to cache tokens or cache information really. An example of this is Elegant Insights. We send on every message, or was doing every message to AI, a full database file of Halo that we've manually done and a massive prompt. It was around 60 to 80,000 tokens. It was costing me around 30 cents a message. Mental, I know.

However, what you used to be able to do is cache that for 5 minutes. So when we built Elegant Insights, we worked out on average, if we have this many users and they're asking this many questions per hour, the caching will reset every 5 minutes and it will on average cost us about this per hour, right?

However, a month before we released Elegant Insights, or two months before we released it, Claude brought out one hour prompt caching, which means those 60 to 80,000 tokens we sent to Claude now sit there for an hour, and when people access them to ask for reports, it costs a fraction of the price. When we built Elegant Insights, it was priced fully to break even. I know, bad at business. However, because Claude, or Anthropic, allowed us to do prompt caching for an hour, I now make around £10 to £15 a month on average if people max their tokens.

So again, prompt caching made a product go from we're breaking even trying to help people to okay, we're actually making a few dollars now. We could invest in a developer if we get x number of users. So that's the difference. Caching is really, really, really important. You get with I think Claude, it's like a 90% saving depending on how you leverage it. So again, massive, massive wins.

Then I want to talk about RAG, or retrieval augmented generation. What the hell is RAG? Well, it's really a fancy name for AI searching your documents. Basically, the way it works is a ticket arrives into your help desk. VPN keeps disconnecting. AI using RAG will then go and search your knowledge base. It will find something similar in there and then it will answer the question, VPN keeps disconnected, or provides a solution based on your documentation.

This is really important because if you don't use something like this to answer questions, you're relying on AI making it up itself or assuming it knows it, which is very dangerous. Or it goes and uses a web search. Now, interestingly, this is like the 15th take of this video. I watched a video last night that Luke from NinjaOne posted about this, where AI is making so much documentation on the web now that AI then searches on. Now, the problem you get here is garbage in garbage out. So if AI is making the documentations with 50% truth, then it goes searches it and makes another documentation that's then 25%, like you can very quickly see how AI searching AI-generated slop can cause a massive problem. And this is the issue we're having in industry right now where there's so much AI-generated stuff out there. If you're using that to reference, is that even correct data? But I won't go off track. The bottom line with RAG is it transforms a helpful AI assistant that you can chat to with something that can go and search through your data.

Now you'll probably see as well, moving on slightly, to vectors and embeddings. It's not quite plain text searching your documentation. If VPN keeps disconnecting comes in as a ticket, it's not searching your database for the words VPN keeps disconnecting. What it's actually using is vector and embeddings, or vector embeddings to be more specific. So what the hell is vectors embedded? If I manage to make it through this take and I actually understand what the hell I'm saying, we might have a good chance at this one.

So this is vector and embeddings, or vector embeddings, is probably better way to articulate it. Vector embeddings. What is it? Well, it's complicated honestly to get your head around. All you really need to know is you send your data to AI and it puts it in a string of numbers, basically an embedding, which looks to me like a GPS coordinate, and the more data you send it, let's say you sent it 15 tickets, each of those tickets would have an embedding score next to it, okay? Then what you're doing when you're trying to match it, it's a game of snap. You go find me things that are like this sentence in my database or whatever, and you're trying to match it based on the numbers, okay? Or this GPS coordinate.

The complicated thing to understand though, and the bit that I always get caught up on trying to explain it, is that an embedding isn't just words put in numbers. It's not just letters put in numbers. It's based contextually. So let's say you put the sentence "I don't know, I went to the park," and then you put the next sentence of "you went to the park." Well, those embedding numbers would look really differently. You would struggle to understand as a human that the word park even exists in there because it will take the context around it and will basically do these embedding scores for you.

Now I think Claude either uses its own embedding or outsources it. I can't remember now how it handles its embedding scores, but in essence, what you need to do is send your data to the AI. You need to handle it or run it through something that's going to vectorise it or embed it in the database so that when AI is searching for data using RAG, it knows what to look for.

So vector and embeddings is really complicated. I think the best way to try and understand it is it's AI's way of organising your data. So let's take my wife for instance. My wife has a system for loading the dishwasher. She has a system for loading the cupboards with food. She has a system for loading the fridge. It means when my wife goes to retrieve things from these places or put things in those places, she knows exactly where to go. She may do it by height or colour or weight or I don't know. I've still not figured it out, but she has a system that allows her to retrieve things quicker.

The problem is when I go to get stuff out of there, I have no idea. However, my shed, my shed, is embedded in a way that I know exactly where everything is. It's got its own place. I know where to search things. I know where to put things because I've built a system. So embeddings really is that for AI, but it does it based on more than just the word dog or the number one, two, three. It based it all around context.

So when we're trying to search for documents or search for solutions or search for, I don't know, matching of things, we embed the question. We then look at that massive embedding score and then we try and find the best fit based on all those numbers and their positions to another embedding to form a game of snap. And we say it must be 50, 60, 70, 80, 90, 95% a match to determine that these two things are the same.

That is as much as I'm going in today with vector and embeddings before my little head falls off. It is a way for AI to search through your data.

Now let's talk about structured outputs. What is a structured output? Well, given it simply, let's talk about HaloPSA. You knew it was coming. We use AI for, let's say, automatic triage. So we send a ticket to AI and we say hello Mr. AI, please, always say please, remember Terminator, please can you give me a category, a priority, a ticket type, and I don't know, impact, urgency, for this ticket. Okay, we ask it those questions and then we need to get some data back.

Now the idea of structured outputs is that we can say give me a category, give me a priority, give me an impact and an urgency rating, and we can say they are required outputs from the AI. Now what AI typically does, or what you want it to do, is return in JSON for you. So we can take a JSON payload from the API and we can take that data and update the ticket.

Now if you don't have a structured output and we just say we want these things, AI might give us that data back in different ways. It might give us it in a different order. It might miss out department because it doesn't think about it, or ticket type, or whatever. You may end up with different outputs. Now unfortunately we're still working with computers here and we always need to expect, or we always expect a certain output from AI in a certain order so we can go and update our Halo, for instance.

So that is why structured outputs are really important. When you start leveraging AI with a tool or a typical application, we need to make sure we are getting a structured response from it. Unless we are doing like email replies, right? If you are asking AI to rewrite some text you've written, we don't need a structured output then because we just need an HTML response that we can put into a ticket. We don't mind AI being a little bit creative with that. So we don't need a structured output. But if we are leveraging something like auto triage, then we 100% want structured outputs.

Finally, on this section of AI jargon, there's a lot more we can cover, but I'm just going to end on this one. Let's talk about hallucinations. What are hallucinations? What does it mean? Am I hallucinating? Are you hallucinating? Who knows?

And I'm going to start by eating a little bit of humble pie. So when AI first hit the scene, I'm a cynic anyway. I think I've worked with cybersecurity for too long. I'm a cynic and was like, "Oh, I don't like the idea of AI doing things because it's a robot." And you know, we can't guarantee the outputs, right? Which is a true statement still. We can't 100% guarantee that if I ask AI the same question 10 times, it's going to reply with the same answer. Doesn't matter what temperature you set.

That is, on this topic, temperature is how creative you want the model to be. So let's say zero is only answer with this thing, and then one is be free as a bird. Make whatever you want. So temperature controls that. But even with like zero temperature, we still can never guarantee that it will always respond in exactly the same way, right? And I was like, that's really dangerous because if we wanted to always do a certain thing, we need to guarantee that, and humans do that.

And then the more I've thought on this topic, the more I realise that it's just a blatant lie. The way I interact with things, pre or post breakfast, or pre or post coffee, can vary. If I'm tired, it varies. If I'm not in the mood, it varies. If I fall out with my kids on the school run and get back in for a meeting, I am frustrated. Therefore I am emotionally involved in those decisions, and then they vary. Whereas actually what we're finding is AI is a lot more predictable than humans typically are.

So when we are categorising tickets, for instance, which was a concern of mine, or concern of mine should I say, I was like oh it might do it different than we've got shitty data. Well I think it's a lot better than humans. I generally think that it's a two-fold thing though. So for me, let's talk about AI triage slightly. We have AI set all the things and then we have a human validate it because AI can hallucinate. It can change its mind. Different models have different outcomes.

AI also does often degrade. So certain models have degradation for a day. So the output is severely lower. They may be, I don't know, oversaturated with people using it. They might have infrastructure problems that can affect the output of AI. I've seen it many times. It is highly annoying. But in essence, I think it's a risk you can't ignore, especially if you're putting AI in front of your end users.

If you're asking it to have conversations, that to me is really dangerous. We're aware of the DPD thing that happened where AI swore at its customers. You've probably seen the Air Canada thing where I think it was, internally, it made up policies for its staff to read or something like that. Like what the hell? Again, probably through misconfiguration, probably through not setting up clear guard rails. However, if you're at the stage like we are, still of learning this, and you don't have 50 hours a week to be on top of AI, you've got to consider the risks. So have a human in the loop there. Have a human validating what it is doing, just sanity checking it for most things.

If it's just generating LinkedIn post ideas for you or writing you a bit of a proposal because it's quicker at it and better, fine. You're still going to validate it anyway. But if you're letting it automatically reply and do things to your clients, then you need to be conscious of hallucinations. And that's what it means, hallucinations. Hallucinations is the fact that it will start to drift in a conversation, if you're running out of context. It will start to make things up on occasion because that's what AI is really, really good at. And it can start to provide unexpected results for you.

So please do be aware of hallucinations and look into the various models you're using and how best to reduce hallucinations based on what output you're trying to get. Remember though, if you're playing with temperature as a way to be don't be creative or be creative, being really strict really gates AI. Having a little bit of creativity, allowing it to think, quote unquote, actually will get you more better outputs on occasion. So it's a dial you need to really tailor in to where you just about want it to sit.

So that's hallucinations. Go away and think about it. And that is all I'm going to talk about in this series about AI jargon. There are loads more, or there's way better qualified people than me online that talk about this in more depth. But I'm giving you what I've learned, or what I think I've learned. If I'm wrong in any of these videos, please do call me out because we're always learning with this. But that is what I think I understand over the past couple of years of AI and the things that you most need to be aware of.

Brilliant. Have a good day. I've been Connor Fagan. Speak to you all soon. Goodbye.

Exceptional experience with Renada Solutions. Both Connor and Robbie are easy to work with, their insightful guidance has been invaluable. Highly recommend their services for any business seeking to unlock the full potential of HaloPSA.
Orbits IT Google Logo

Our Core Services

Offering support to enable sustainable success for your organisation.

Consultation Harness the transformative potential of an agnostic advice tailored to your unique business needs. From PSA implementation to ongoing support, our exceptional consultation services pave the way for extraordinary success. Find out more
Virtual Admin Let us handle the technical heavy lifting. Our expert team builds solutions, creates powerful reports and dashboards, and develops automated integrations - giving you more time to focus on what matters most: your clients. Find out more
Product Onboarding We understand that the first steps in adopting a new product can be daunting, we are here to guide you through every stage of the process with precision and clarity. From initial setup to advanced features, maximise the value of your product from day one. Find out more
Virtual Chief Technology Officer (vCTO) Benefit from a remote and adaptable technology expert to seamlessly combine strategic guidance and effective leadership to propel your business to new heights and empower your organisation’s technology ability. Find out more
Where to next? Get the cutting-edge tools to support your MSP business. Contact us today to receive a bespoke quote tailored to your specific needs.