Painting of a ship sailing across blue seas
← Back to blog

TOKEN LIFECYCLE

OAuth Refresh Tokens: Fixing invalid_grant

One error code, eleven causes. A decision table for invalid_grant, what rotation does to concurrent refreshes, and the real expiry numbers for Google, Entra, Slack, Salesforce and GitHub.

•Nov 24, 2025•Updated Sep 29, 2026•16 min
OAuthTokensArchitecture

TL;DR

invalid_grant is not one bug. It is a single error code standing in for about eleven unrelated failures, and the only way to fix it is to work out which one you have before you change any code.

RFC 6749 wrote the ambiguity into the spec. The definition covers a grant that is "invalid, expired, revoked, does not match the redirection URI used in the authorization request, or was issued to another client", all returned as HTTP 400 with the same three-letter string.

The common reading is that the user revoked your app. Sometimes they did. More often it was a seven-day expiry on a Google project still in Testing, a 90-day inactivity window at Microsoft, a clock two minutes off on the box signing a JWT assertion, or your own retry loop replaying a refresh token that rotation had already killed.

We are not claiming rotation is the problem. Rotation is the right default and both the Security BCP and OAuth 2.1 require it or sender-constraining for public clients. It just turns one class of concurrency bug into an outage, and almost nobody writes the lock that prevents it.

What to do today: log the error_description alongside the code, persist the new refresh token in the same transaction that consumes the old one, put a per-connection lock around the refresh, and never retry a 400.

Overview

Take a reconciliation job that runs at 02:00 against a connected accounting account, plus the same integration installed for 400 other tenants. One morning eleven of them fail with {"error":"invalid_grant"} and the rest are fine. Nothing shipped. Nothing in the provider's status page.

That is the shape of the problem. The refresh exchange is one HTTP request and roughly nobody gets it wrong. The diagnosis is what costs the day, because the code carries no information about which of a dozen causes fired and the causes have opposite fixes. Re-consent fixes a revocation and does nothing for clock skew. Retrying fixes a network blip and actively destroys a rotated token.

The exchange that produces the token pair is covered separately in Slack OAuth in Python and storage in how to store OAuth tokens securely. This post is about the part after that: what the error actually means, how to tell the causes apart from the response you already have, and what the five providers most people integrate really do.

Treat invalid_grant as a category, not a diagnosis. Until you can name which cause fired, any fix you apply is a guess, and one of the popular guesses makes it worse.

The Refresh Flow, and Where It Can Fail

Getting a refresh token starts at consent: request offline_access, or whatever the provider calls its equivalent, during the authorization step. The token response then carries access_token, refresh_token, expires_in and token_type.

When the access token expires, you exchange the refresh token at the token endpoint:

http
POST /oauth/token HTTP/1.1
Host: provider.example.com
Content-Type: application/x-www-form-urlencoded

grant_type=refresh_token&
client_id=YOUR_CLIENT_ID&
client_secret=YOUR_CLIENT_SECRET&
refresh_token=REFRESH_TOKEN

RFC 6749 section 6 says the server MUST require client authentication for confidential clients, MUST ensure the refresh token was issued to that authenticated client, and MUST validate the token. Any of those three checks failing produces an error response under section 5.2, which specifies "an HTTP 400 (Bad Request) status code (unless specified otherwise)".

Provider token endpointYour backendProvider token endpointYour backendaccess token expiresPOST grant_type=refresh_token200 + new access token (+ new refresh token, if rotating)persist, then resume API calls

Figure 1 — The happy path. The only interesting line is the last one, because persisting the response is where rotation bugs live.

Section 5.2 reserves a different code for a different failure. invalid_client means "Client authentication failed (e.g., unknown client, no client authentication included, or unsupported authentication method)", and the server MAY answer 401 for it. Plenty of providers collapse that into invalid_grant anyway, which is why the first thing to check after a mass failure is whether anyone rotated the client secret.

Diagnosing invalid_grant

Here is the table we would keep next to the runbook. The middle column is what the response body alone tells you; the right column is the signal that actually separates the cause from its neighbours.

CauseWhat the response looks likeHow you tell it apart
Rotated token replayed400 invalid_grant, immediately, on a token you believe is currentTwo refreshes for the same connection within milliseconds in your logs, usually from two workers. Under OAuth 2.1 the whole grant dies, not just the token
User revoked the app400 invalid_grant; Entra returns AADSTS50173, Google's description says the token was expired or revokedAffects one user across every one of your tokens for them, at a time that matches nothing you did
Password change or resetEntra AADSTS50173: "The user might have changed or reset their password"Entra only revokes password-based tokens on a user-initiated change; confidential-client tokens "stay alive". Google revokes only if the grant includes Gmail scopes
Seven-day expiry, Google Testing app400 invalid_grant exactly seven days after each consent, for everyoneYour OAuth consent screen is external user type with publishing status Testing. Flip it to In production
Inactivity expiryEntra AADSTS70008 or AADSTS700082, both "The refresh token has expired due to inactivity"Compare against your last-successful-refresh timestamp. Google's window is six months, Entra's MaxInactiveTime is 90 days
100-token-per-user cap, Google400 invalid_grant on an old token while a newer one for the same user worksCount how many times that user re-consented. Google "automatically invalidates the oldest refresh token without warning"
Clock skew on a JWT assertioninvalid_grant with "Invalid JWT: Token must be a short-lived token (60 minutes) and in a reasonable timeframe."The literal message. Google also fires it when exp is more than 65 minutes past iat, or below it
Wrong client credentialsShould be invalid_client, often arrives as invalid_grantFails for 100% of connections at once, starting at a deploy or a secret rotation
Wrong redirect_uriinvalid_grant; RFC 6749 names this case explicitlyOnly ever on the authorization-code exchange. A refresh request sends no redirect_uri, so this cannot be your cause at 02:00
Sandbox against productionSalesforce invalid_grant for every callThe token was minted at test.salesforce.com or a --sandbox.my.salesforce.com host and you are posting to login.salesforce.com. Fails 100% in one environment, 0% in the other
Authorization code already spentinvalid_grant, or a provider-specific code such as Slack's invalid_codeInstall-time only. Slack's temporary code "expires after ten minutes" and is single use

Table 1 — Eleven causes behind one error string, and the signal that identifies each.

Two columns of that table are worth turning into instrumentation: blast radius (one user, one tenant, or everybody) and timing relative to the last successful refresh. Those two numbers separate most of the rows without reading a provider doc, and neither is in the error payload, so you have to record them yourself.

The other habit worth building is logging error_description verbatim. Google, Microsoft and Salesforce all put the actual cause there while keeping error at invalid_grant. Salesforce's authorization error list folds "Invalid authorization code", "Invalid user credentials", "Invalid assertion", "Invalid audience", "IP restricted or invalid login hours", "User hasn't approved the connected app" and "For the refresh token flow, the refresh or access token is expired" all under that one code. Throwing the description away and keeping only the code is the single most expensive line of error handling we see in this space.

Rotation and the Replay Problem

RFC 6749 section 6 left rotation optional: "The authorization server MAY issue a new refresh token, in which case the client MUST discard the old refresh token and replace it with the new refresh token." The MUST is on the client. Section 10.4 then describes why a server would bother, and the sentence is still the clearest statement of the idea: "If a refresh token is compromised and subsequently used by both the attacker and the legitimate client, one of them will present an invalidated refresh token, which will inform the authorization server of the breach."

RFC 9700, the OAuth 2.0 Security Best Current Practice published in January 2025, hardened that into a requirement. Authorization servers "MUST utilize one of these methods to detect refresh token replay by malicious actors for public clients": sender-constrained refresh tokens via mTLS or DPoP, or refresh token rotation. The consequence it spells out is the part people miss. The server "cannot determine which party submitted the invalid refresh token, but it will revoke the active refresh token." OAuth 2.1, currently at draft-ietf-oauth-v2-1-16 and dated September 2026, goes one step further: it revokes "the active refresh token as well as the access authorization grant associated with it."

Read that as an operational fact rather than a security one. A replay does not fail the one request. It kills the connection and sends your user back through consent.

Rotation converts a silent theft into a loud outage, and it cannot tell your bug from an attacker. A second worker refreshing the same connection presents exactly the evidence the spec says to treat as a breach.

This is where the concurrency problem bites. Two processes hold the same refresh token. Both notice the access token is stale. Both POST.

ProviderWorker 2Worker 1ProviderWorker 2Worker 1grant revoked, RT2 now dead toorefresh(RT1)200, issues RT2, invalidates RT1refresh(RT1)400 invalid_grant (replay)API call with new access token401

Figure 2 — The concurrent-refresh race. Worker 2 is not an attacker, but the authorization server has no way to know that, and under OAuth 2.1 it revokes the grant.

The fix is single-flight: exactly one refresh in progress per connection, everybody else waits for its result. In practice that means taking a lock keyed on the connection before you decide to refresh, re-reading the stored token after you acquire it (someone may have refreshed while you queued), and releasing only after the new token is committed.

python
async def get_access_token(conn_id: str) -> str:
    row = await store.load(conn_id)
    if row.expires_at - time.time() > SKEW_SECONDS:
        return row.access_token

    async with locks.acquire(f"refresh:{conn_id}", ttl=30):
        row = await store.load(conn_id)          # re-read: someone may have won
        if row.expires_at - time.time() > SKEW_SECONDS:
            return row.access_token
        resp = await provider.refresh(row.refresh_token)
        await store.commit(conn_id, resp)        # both tokens, one transaction
        return resp["access_token"]

A Postgres advisory lock, a Redis lock with a TTL, or a SELECT ... FOR UPDATE on the connection row all work. An in-process mutex does not, once you run more than one process, and that is the version most codebases ship with.

Set SKEW_SECONDS generously, 60 to 300 seconds. Refreshing early is nearly free. Discovering expiry via a 401 halfway through a batch is not.

What the Providers Actually Do

Vendor behaviour diverges more than the spec suggests, and the numbers are the useful part.

ProviderAccess token lifeRefresh token lifeRotation
Google1 hour typicalSix months of inactivity; 7 days if the consent screen is external and in Testing; oldest dropped past 100 per user per client IDNot by default for web server apps
Microsoft EntraRandom 60 to 90 minutes, average 75; 24 to 28 hours with Continuous Access Evaluation90 days for most scenarios, 24 hours for single-page apps and email one-time passcode; MaxInactiveTime 90 days, max age "Until-revoked"Yes, but the old token is not invalidated
Slack (rotation enabled)12 hours, expires_in 43200Single use, revoked "after a short grace period"Yes, mandatory once enabled
SalesforceSession-length dependentGoverned by the connected app's refresh token settingOptional, via Enable Refresh Token Rotation
GitHub App user tokens8 hours, expires_in 288006 months, refresh_token_expires_in 15897600Yes, when expiration is enabled

Table 2 — Documented lifetimes, read from each vendor's own reference pages.

Three of those rows have a trap in them.

Microsoft's is the most surprising. Entra rotates, but its refresh token documentation says plainly: "Refresh tokens replace themselves with a fresh token upon every use. The Microsoft identity platform doesn't revoke old refresh tokens when used to fetch new access tokens." So the replay detection that rotation exists to provide is not what you get here, and the burden shifts entirely to you: "Securely delete the old refresh token after acquiring a new one." Entra also stopped letting you configure any of this. As of 30 January 2021 refresh and session token lifetimes are no longer settable through token lifetime policies, and existing properties are ignored.

Slack's is the sharpest. Once an app opts into token rotation, the access token expires every 12 hours and the refresh token is genuinely single-use. The docs also warn against refreshing in a tight loop: "If you refresh your credentials repeatedly before expiration...we will enforce a limit of 2 active tokens." Migrating an existing app uses oauth.v2.exchange, and it is one-way per token: "You won't be able to exchange the same access token for a refresh token more than once."

Google's is the one that generates the most search traffic. A project whose OAuth consent screen is configured for external users and left in Testing "is issued a refresh token expiring in 7 days". That is not a bug in your code and no amount of refresh logic survives it. The other Google cliff is the cap: "There is currently a limit of 100 refresh tokens per Google Account per OAuth 2.0 client ID", and past it "creating a new refresh token automatically invalidates the oldest refresh token without warning". If you mint a token per device or per re-consent and never clean up, you will eventually break your own oldest connection for a user who did nothing.

Salesforce revokes hard on reuse when rotation is on: reusing an invalidated refresh token revokes the current refresh token and all associated access tokens. Auth0 goes further, denying "all subsequent requests...until the user re-authenticates" and invalidating the whole token family.

Storage, Retries, and the Loop That Burns the Grant

On a successful refresh, persist four things before you use the new access token for anything: the access token, the refresh token from the response if one was returned, an absolute expiry computed from expires_in at receipt time, and the timestamp of the refresh itself. The last one is not decoration. It is the field that tells you, six weeks later, whether row five of Table 1 applies.

Two ordering rules matter more than the storage technology.

First, commit the new refresh token in the same transaction that marks the old one spent. If your process dies between "provider issued RT2" and "database has RT2", the connection is gone and no retry recovers it. Write first, use second.

Second, if the provider returns a response with no refresh_token field, keep the one you have. That is RFC 6749 behaviour, not an error, and overwriting your stored token with None is a popular way to manufacture an outage. GitHub is explicit that when you disable token expiration the refresh_token and expires_in fields are simply omitted.

Retry behaviour is where an unconditional loop does real damage. A 400 from the token endpoint is a verdict, not a transient failure. Retrying it against a rotating provider replays a token the server has already invalidated, which is the exact evidence RFC 9700 tells the server to treat as a breach.

python
resp = await http.post(TOKEN_URL, data=payload)

if resp.status_code == 400:
    body = resp.json()
    await store.mark_broken(conn_id, body.get("error"),
                            body.get("error_description"))
    raise ReconsentRequired(body.get("error_description"))

if resp.status_code in (429, 500, 502, 503, 504):
    raise Retryable(retry_after=resp.headers.get("Retry-After"))

resp.raise_for_status()

Retry 429 and 5xx with backoff. Never retry 400. And be careful with generic HTTP clients that retry on connection errors, because a refresh that succeeded at the provider but timed out on your side is the worst case: the server has rotated, you have not, and your automatic retry sends the dead token.

The rule we would enforce in code review: the only two outcomes of a refresh are "new tokens committed" and "this connection needs consent again". Anything in between is a retry loop you will regret.

Watch the ratio of invalid_grant to successful refreshes per provider per hour, not the absolute count. One connection failing is a support ticket. Five percent of a tenant failing inside ten minutes is a provider-side change, and those want different responses.

Where Fabriq Fits

Agentic Fabriq holds credentials per user in a vault and attaches them server-side at the moment of the call, so the agent making the request never handles the refresh token. Refresh runs automatically behind that boundary with last-used tracking on each connection, which is what makes the inactivity rows in Table 1 diagnosable rather than mysterious.

What it does not do is reach into the provider for you. Clearing a vault copy is not the same act as invalidating a grant upstream, so the call to https://oauth2.googleapis.com/revoke or the provider's equivalent is still yours to make.

Frequently Asked Questions

Why does my refresh token expire after 7 days? Almost certainly because your Google Cloud project's OAuth consent screen is set to external user type with a publishing status of Testing. Google issues those projects "a refresh token expiring in 7 days". Publishing the app to In production stops it. Nothing in your refresh code is wrong.

What does "invalid_grant: Token has been expired or revoked" mean? Google's troubleshooting guidance describes it as a token from the Google Authorization Server that "has either expired or has been revoked", and points at the refresh token expiration rules. In practice it is the seven-day Testing expiry, six months of inactivity, a user revoking access, or the 100-token cap evicting your oldest token. Check which of those fits before assuming revocation.

Can I use the same refresh token twice? Only if the provider does not rotate. Where rotation is on, the second use is treated as replay: RFC 9700 says the server "will revoke the active refresh token", and OAuth 2.1 revokes the associated grant too. Slack's rotating tokens are "designed to be used once". Microsoft Entra is the exception that proves the point, since it rotates but "doesn't revoke old refresh tokens when used to fetch new access tokens".

Why does refreshing fail only when my job runs in parallel? Because two workers are racing on the same refresh token and one of them presents an already-rotated value. Put a lock keyed on the connection around the refresh, re-read the stored token after acquiring it, and return the winner's token to the losers.

Do refresh tokens expire if I never use them? Usually. RFC 9700 says refresh tokens "SHOULD expire if the client has been inactive for some time". Google's window is six months. Entra's MaxInactiveTime is 90 days and surfaces as AADSTS70008 or AADSTS700082. A keep-alive refresh on a schedule well inside the window is a legitimate defence.

Does a password change kill my refresh token? It depends on the token class. Entra revokes password-based tokens when the user changes their own password, while confidential-client tokens "stay alive"; a user or admin revoking refresh tokens outright kills every class. Google revokes on password change only when the grant includes Gmail scopes.

Is invalid_grant ever my client secret? It should be invalid_client under RFC 6749, which also allows a 401 for that case, but several providers answer invalid_grant regardless. The tell is blast radius: a bad secret fails every connection at once, and a dead grant fails one.

Should I retry an invalid_grant? No. It is a 400, which means a verdict. Mark the connection as needing consent, alert, and stop. Retrying against a rotating provider replays an invalidated token and can take down the grant that the first failure had not yet touched.

How do I stop hitting Google's 100-token limit? Stop minting a new token on every consent you do not need, and delete tokens you have stopped using rather than leaving them to be evicted. The cap is per Google Account per client ID, and eviction is silent.

Conclusion

The refresh exchange is one request and it is not where the engineering is. The engineering is in accepting that invalid_grant is a category label, recording enough context at failure time to resolve it to a cause, and building a refresh path that can only end in two states.

Do three things and most of this disappears. Log error_description next to error, because the providers put the answer there. Hold a per-connection lock across the refresh and commit both tokens in one transaction, because rotation punishes concurrency harder than it punishes attackers. And treat a 400 as terminal, because the retry is what turns one broken connection into a revoked grant.

A refresh token that has never failed is not one that is safe. It is one nobody has tested. Force a failure in staging, revoke a grant by hand, and watch what your code does. Whatever it does then is what it will do at 02:00.

Sources

All URLs read 2026-09-29.