Tech
Gemini free limits change from 9 October: how to save
From 9 October, reports say the free Gemini allowance will depend on the computing demands of tasks, and model choice will narrow. We explain what is changing and how to write more economical prompts.
2026-10-08 · 5 min read
From October, Google will count usage differently for people using the Gemini artificial intelligence service with a free account. According to HVG's report of 7 October, from 9 October 2026 the allowance for free users will be determined not by a pre-set number of messages, but by the computing capacity required by tasks. The paper also says that only the Flash-Lite model will be available with free accounts. That means how someone phrases their questions will also matter. Below, we explain what is changing and how to write requests that require fewer resources.
What is changing for free accounts from 9 October?
According to HVG's report, the change has two main elements. One is that usage limits will be aligned with the computing demands of tasks. The other is that free accounts will be able to access only the least resource-intensive Flash-Lite model, while the more capable Flash and Pro models will be tied to paid plans.
ZDNET reported the change in similar terms: according to the outlet, free users will be restricted to Gemini's weakest model, while AI Plus and AI Pro subscriptions offer stronger models.
Based on the reports, it is not worth expecting a fixed daily message count. If the allowance depends on the computing demands of requests, then the same number of messages will probably last longer for a user asking short questions than for someone requesting long analyses or the processing of lengthy documents.
Computing demand instead of message count: what is the difference?
A message-count-based limit was simple: every request sent counted as one unit, whether it was a one-word translation or an analysis of a report several pages long. A system based on computing capacity, by contrast, takes into account how much work the model running in the background performs when generating each answer.
Artificial intelligence systems such as Gemini process text by breaking it into smaller units known as tokens, in other words word fragments. The pasted text, the question and the answer all consist of tokens, and the more tokens that have to be processed or generated, the greater the computing load.
What Google treats as a resource is illustrated well by the Gemini API developer documentation. According to this, in developer access the limits are governed by requests per minute (RPM), tokens per minute (TPM) and requests per day (RPD), while thinking tokens and context size also count towards usage. This is the rule set for developers, not for the everyday Gemini app interface, but the logic behind it is similar: longer and more complex work uses more capacity.
What does effort level mean, and why does it matter?
According to ZDNET's report, so-called effort levels will appear for requests. These indicate how much internal “thinking” the system does before responding: at a lower level it answers faster, with fewer intermediate steps, while at a higher level it works more thoroughly.
Since Google's developer documentation says that thinking tokens also count towards the allowance, a higher effort level is likely to consume free capacity more quickly. For a spelling correction, a short translation or summarising a paragraph, deep thinking is generally not needed. The higher level may be justified mainly for complex logical, mathematical or multi-step tasks.
Where is the allowance wasted unnecessarily?
According to Google Cloud's guide on context and instruction settings, unnecessary context increases the number of tokens used, while structured and concise instructions reduce the computing load. According to an analysis by Inventive HQ, wordy, repetitive instructions can use 50–80 per cent more resources than necessary. This estimate primarily relates to developer token costs, so it is best treated as an indication of scale.
The most common wasteful habits:
- Long introductory preambles: polite formulas and explanatory padding around the question also count as tokens.
- Repeated instructions: if someone describes the same thing in several different ways, the answer will not be more accurate, only the request will be longer.
- Pasting full documents: if the issue concerns only one paragraph, there is no need to insert the entire text.
- Open-ended questions with no format constraints: these often result in long, general answers.
- Conversations that drag on indefinitely: because context size also matters, continuing a long thread is likely to require more capacity than starting a new conversation.
Wasteful and economical prompts: two examples
First example, wordy version: “Hi! I hope you're well. I have a question because I'd like to buy a used washing machine, and I do not really know much about them. Could you write to me about what is worth paying attention to, preferably in detail, but not too long, and if possible please cover every important point?”
Concise version: “List five things to consider when buying a used washing machine, with one sentence for each point.” It asks for the same information, but the length and format are clear, so the answer will also be shorter.
Second example, wasteful version: pasting an entire long email exchange with the question, “What do you think about this?”
Economical version: “Summarise the main point of the email below in three bullet points, and state what deadline it includes:” — followed only by the relevant message. According to Inventive HQ's analysis, it is precisely this kind of compression and clear constraints on the response format that reduce token consumption.
Practical rules for economical use
- Start with the task: it should be clear from the first sentence what Gemini needs to do.
- Specify the format and length: for example, “in three bullet points”, “in no more than five sentences” or “in a table”.
- Paste only the necessary text: according to Google's guide, unnecessary background also increases the load.
- Choose a lower thinking level for a simple task, if the interface allows it.
- Open a new conversation for a new topic, so that earlier history does not increase the context.
- Be specific when asking for corrections: if the answer is not suitable, ask for a targeted change rather than regenerating the whole thing.
Who should consider a subscription?
According to ZDNET, the stronger models will remain available in the AI Plus and AI Pro packages. Based on a report by the British trade publication Computing UK, discounted AI Pro access for higher education students will remain in place after the changes, so university students may want to check whether they are eligible.
For someone who uses Gemini occasionally for short questions, more deliberate phrasing will probably be enough to make the free allowance last. Anyone who regularly works with long documents, program code or complex analyses may want to compare the paid packages with their own needs. Before deciding, it is sensible to check Google's current terms, because limits may vary depending on the nature of the requests.
In short: from 9 October, the free Gemini allowance will be consumed not by the number of messages, but by the weight of the requests. Anyone who gives short, clear tasks, pastes only the necessary text, and does not ask for deep thinking on simple questions can get more out of the same allowance.
Sources used
- 1.Gemini API Rate Limits & Pricingai.google.devverified
- 2.Context and Prompt Configurationdocs.cloud.google.comverified
- 3.Use Gemini for free? You'll soon be limited to its weakest AI modelzdnet.comverified
- 4.Google to restrict Gemini model access for free userscomputing.co.ukverified
- 5.Optimizing Prompts to Reduce Token Usage and Costsinventivehq.comverified
These sources were used during our editorial fact check.