> ## Documentation Index
> Fetch the complete documentation index at: https://docs.dodopayments.com/llms.txt
> Use this file to discover all available pages before exploring further.

# OpenAI Billing Model

> Deconstruct OpenAI's token-based billing with prepaid fiat credits, then rebuild it with Dodo Payments credit entitlements, usage meters, and rate tiers.

OpenAI combines prepaid fiat credits for API usage with flat-rate subscriptions for its consumer products. The prepaid side gives OpenAI cash up front, and developers can scale usage without a sales conversation. Many AI companies copy this model.

## Why OpenAI's Model is the Standard

Traditional SaaS billing doesn't handle the variable cost of AI usage well. OpenAI's model solves three problems at once:

1. **Predictable Revenue and Low Risk**: Because API usage is prepaid, users can't run up bills they can't pay. OpenAI receives the money up front, and the user spends it as they use the service.
2. **Scalability for Developers**: A \$5 top-up is a low barrier to entry. As an application grows, developers can automate top-ups or buy larger packs. Starting is cheap, and usage can grow without a plan change.
3. **User Psychology**: Credits denominated in US dollars, instead of abstract "tokens" or "points", make the value clear. The balance works like a prepaid account for AI services, which makes budgeting easier for companies.

## How OpenAI Bills

OpenAI runs two billing models for different users.

1. **API (Pay-as-you-go)**: The API uses prepaid, dollar-denominated credits. Users top up their accounts with \$5, \$10, \$50, or more. The credits show a dollar value but can't be used outside OpenAI. OpenAI bills per token, with different rates for input and output tokens. Purchased credits expire one year after purchase and are non-refundable ([OpenAI Help Center](https://help.openai.com/en/articles/8264644-setting-up-and-managing-prepaid-api-billing)). When a balance reaches \$0, API calls fail.
2. **ChatGPT Plus, Business, and Enterprise**: These are flat-rate subscriptions. ChatGPT Plus costs \$20 per month, and the Business plan (formerly Team) costs \$25 per user per month when billed monthly. They have soft usage caps: heavy users move to a smaller model instead of being blocked.
3. **Spend-based rate tiers**: As total spend grows over time, the account unlocks higher API rate limits. Access grows with billing history.

The token prices below are the published rates at the time of writing:

| Model | Pricing | Input Tokens | Output Tokens |
| :- | :- | :- | :- |
| GPT-4o | Usage-based | \$2.50 / 1M | \$10.00 / 1M |
| GPT-4o-mini | Usage-based | \$0.15 / 1M | \$0.60 / 1M |
| o1 | Usage-based | \$15.00 / 1M | \$60.00 / 1M |

| Plan | Price | Type |
| :- | :- | :- |
| Free | \$0 | Limited access |
| Plus | \$20 / mo | Subscription with soft caps |
| Business | \$25 / user / mo | Per-seat subscription |
| Enterprise | Custom | Invoiced billing |

## What Makes It Unique

Four characteristics make OpenAI's billing effective for AI services:

* **Fiat-denominated credits**: Credits are in US dollars, so they feel like money. Developers can read the price of a request directly.
* **Long expiry**: Purchased credits last a year, which reduces "use it or lose it" pressure. Users are comfortable topping up larger amounts.
* **Multi-dimensional metering**: Input and output tokens are tracked separately but deduct from the same balance. OpenAI can price expensive output tokens higher than input tokens.
* **Trust tiers**: Rate limits that rise with total spend reward long-term customers and encourage them to stay.

## Strategic Advantages

The model reinforces itself. Low entry costs bring in developers. Prepaid credits provide immediate cash flow. Usage-based pricing means OpenAI earns more as developers succeed. Subscriptions add a steady baseline of revenue from non-developers.

## Build This with Dodo Payments

You can build OpenAI's billing model with Dodo Payments. Use Credit-Based Billing for the API side and standard subscriptions for the ChatGPT Plus side.

<Steps>
  <Step title="Create a Fiat Credit Entitlement">
    In your Dodo Payments dashboard, go to **Products → Credits** and click **Create Credit**. This credit is the central balance for each user.

    * **Credit Type:** Fiat Credits, with **Unit Currency** set to USD
    * **Credit Expiry:** Custom, 365 days (matches OpenAI's one-year expiry), or Never
    * **Rollover:** Not needed (credits don't reset each cycle)
    * **Allow Overage:** Disabled

    Fiat credits use two decimal places, so one credit is one dollar and balances track cents. Dodo Payments doesn't block usage when a balance reaches zero. To make API calls fail at \$0 like OpenAI, check the balance in your application before each request (see **Handle Balance Depletion** below).
  </Step>

  <Step title="Create Top-Up Products">
    Create one-time payment products for different credit packs, such as \$5, \$10, \$50, and \$100. Attach your fiat credit to each product.

    Set the number of credits issued to the pack's dollar value. A \$50 pack issues 50 credits.

    ```typescript theme={null}
    import DodoPayments from 'dodopayments';

    const client = new DodoPayments({
      bearerToken: process.env.DODO_PAYMENTS_API_KEY,
    });

    const session = await client.checkoutSessions.create({
      product_cart: [
        { product_id: 'pdt_credit_pack_50', quantity: 1 }
      ],
      customer: { email: 'developer@example.com' },
      return_url: 'https://yourapp.com/dashboard'
    });
    ```
  </Step>

  <Step title="Create Usage Meters">
    Create two meters to track token usage:

    * `llm.input_tokens`: Sum aggregation on the `tokens` property.
    * `llm.output_tokens`: Sum aggregation on the `tokens` property.

    On your usage-based product, toggle **Bill usage in Credits** on both meters and select the fiat credit. Then set **Meter units per credit** for each.

    ### Calculating Meter Units per Credit

    To match OpenAI's GPT-4o pricing, work out how many tokens cost \$1, which is one fiat credit:

    * **Input Tokens:** 1,000,000 tokens / \$2.50 = 400,000 tokens per \$1.
    * **Output Tokens:** 1,000,000 tokens / \$10.00 = 100,000 tokens per \$1.

    In the Dodo Payments dashboard, set **Meter units per credit** to 400,000 for input and 100,000 for output. Dodo Payments divides each meter's aggregated tokens by this value to get the credits to deduct.
  </Step>

  <Step title="Send Usage Events">
    After each LLM request, send the usage to Dodo Payments. One request can carry both the input and the output event. This snippet reuses the `client` from the previous step.

    ```typescript theme={null}
    // requestId and customerId come from your application's request context.
    await client.usageEvents.ingest({
      events: [{
        event_id: `req_${requestId}`,
        customer_id: customerId,
        event_name: 'llm.input_tokens',
        timestamp: new Date().toISOString(),
        metadata: {
          model: 'gpt-4o',
          tokens: 1500
        }
      }, {
        event_id: `req_${requestId}_out`,
        customer_id: customerId,
        event_name: 'llm.output_tokens',
        timestamp: new Date().toISOString(),
        metadata: {
          model: 'gpt-4o',
          tokens: 800
        }
      }]
    });
    ```
  </Step>

  <Step title="Handle Balance Depletion">
    Check the user's balance before you process an API request. If the balance is zero or negative, reject the request, for example with a `402` status.

    ```typescript theme={null}
    async function checkCreditsBeforeRequest(customerId: string) {
      const balance = await client.creditEntitlements.balances.retrieve(customerId, {
        credit_entitlement_id: 'cde_your_credit_entitlement_id',
      });

      if (Number(balance.balance) <= 0) {
        throw new Error('Insufficient credits. Please top up your account.');
      }
    }
    ```

    ### Handling Low Balance Webhooks

    Notify users before they reach \$0. Set a **Low Balance Threshold** when you attach the credit, then send an email or in-app notification when the `credit.balance_low` webhook arrives.

    ```typescript theme={null}
    import DodoPayments from 'dodopayments';
    import express from 'express';

    const app = express();
    app.use(express.raw({ type: 'application/json' }));

    const client = new DodoPayments({
      bearerToken: process.env.DODO_PAYMENTS_API_KEY,
      webhookKey: process.env.DODO_PAYMENTS_WEBHOOK_KEY,
    });

    // sendLowBalanceEmail is your application's own function.
    app.post('/webhooks/dodo', async (req, res) => {
      try {
        const event = client.webhooks.unwrap(req.body.toString(), {
          headers: {
            'webhook-id': req.headers['webhook-id'] as string,
            'webhook-signature': req.headers['webhook-signature'] as string,
            'webhook-timestamp': req.headers['webhook-timestamp'] as string,
          },
        });

        if (event.type === 'credit.balance_low') {
          const { customer_id, available_balance } = event.data;
          await sendLowBalanceEmail(customer_id, available_balance);
        }

        res.json({ received: true });
      } catch (error) {
        res.status(401).json({ error: 'Invalid signature' });
      }
    });
    ```

    <Tip>
      OpenAI offers auto recharge, which buys more credits when the balance falls below a threshold the user sets.
    </Tip>
  </Step>

  <Step title="Build the ChatGPT Subscription Side (Optional)">
    To offer a subscription plan like ChatGPT Plus, create a separate subscription product in Dodo Payments. It doesn't need a credit entitlement.

    For a Team plan, use [seat-based billing](/features/seat-based-billing): a per-seat add-on whose quantity is the number of users.

    ```typescript theme={null}
    const session = await client.checkoutSessions.create({
      product_cart: [
        { product_id: 'pdt_plus_subscription', quantity: 1 }
      ],
      customer: { email: 'user@example.com' },
      return_url: 'https://yourapp.com/billing'
    });
    ```

    ### Implementing Soft Caps

    To build soft caps, track subscription users' usage with the same meters but without linking them to a credit. In your application, check the usage for the current billing period.

    ```typescript theme={null}
    // getUsageForCurrentPeriod and SOFT_CAP_THRESHOLD are defined in your application.
    async function checkSubscriptionUsage(customerId: string) {
      const usage = await getUsageForCurrentPeriod(customerId);
      
      if (usage > SOFT_CAP_THRESHOLD) {
        // Route to a smaller model instead of blocking
        return 'gpt-4o-mini';
      }
      
      return 'gpt-4o';
    }
    ```
  </Step>
</Steps>

## Accelerate with the LLM Ingestion Blueprint

The steps above build and send usage events by hand. The [LLM Ingestion Blueprint](/developer-resources/ingestion-blueprints/llm) instead wraps your OpenAI client and tracks tokens automatically.

```bash theme={null}
npm install @dodopayments/ingestion-blueprints
```

```typescript theme={null}
import { createLLMTracker } from '@dodopayments/ingestion-blueprints';
import OpenAI from 'openai';

const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });

const tracker = createLLMTracker({
  apiKey: process.env.DODO_PAYMENTS_API_KEY,
  environment: 'live_mode',
  eventName: 'llm.chat_completion',
});

// customerId and prompt come from your application.
const trackedClient = tracker.wrap({
  client: openai,
  customerId: customerId,
});

// Every API call now automatically tracks token usage
const response = await trackedClient.chat.completions.create({
  model: 'gpt-4o',
  messages: [{ role: 'user', content: prompt }],
});

// inputTokens, outputTokens, and totalTokens are sent automatically
console.log('Tokens used:', response.usage);
```

The blueprint reads `inputTokens`, `outputTokens`, and `totalTokens` from every API response and sends them, with `model`, as event metadata. Set your meter's **Over Property** to the token key you want to bill.

<Tip>
  The LLM Blueprint supports OpenAI, Anthropic, Groq, Google Gemini, OpenRouter, and the Vercel AI SDK. See the [full blueprint documentation](/developer-resources/ingestion-blueprints/llm) for provider-specific examples and advanced configuration.
</Tip>

## Implementing Spend-Based Rate Tiers

OpenAI's rate tiers manage capacity by trust. To build them, track each customer's lifetime spend.

1. **Track Lifetime Spend:** Listen for `payment.succeeded` webhooks and add the payment amount to a `total_spend` field for that customer in your database. Amounts are in the smallest currency unit, so 5000 is \$50.00.
2. **Define Tiers:** Map spend amounts to rate limits:
   * Tier 1: \$0 - \$50 spend -> 3 RPM
   * Tier 2: \$50 - \$250 spend -> 10 RPM
   * Tier 3: \$250+ spend -> 50 RPM
3. **Enforce Limits:** In your API middleware, look up the customer's tier and apply its rate limit.

```typescript theme={null}
// db and the TIER_*_LIMITS constants are defined in your application.
async function getRateLimitForCustomer(customerId: string) {
  const customer = await db.customers.findUnique({ where: { id: customerId } });
  const totalSpend = customer.total_spend;

  if (totalSpend >= 25000) return TIER_3_LIMITS; // $250.00
  if (totalSpend >= 5000) return TIER_2_LIMITS;  // $50.00
  return TIER_1_LIMITS;
}
```

## Full Implementation Example: The API Proxy

In production, an API proxy usually sits between your users and the LLM provider. The proxy authenticates the request, checks credits, and reports usage.

```mermaid theme={null}
sequenceDiagram
    participant U as User
    participant P as Your API Proxy
    participant D as Dodo Payments
    participant O as OpenAI
    U->>P: API Request
    P->>D: Check Credit Balance
    D-->>P: Balance Available
    P->>O: Forward Request
    O-->>P: Response + Token Count
    P->>D: Ingest Usage Events
    D-->>P: Events Accepted
    P-->>U: Return Response
```

The handler below implements the proxy:

```typescript theme={null}
import DodoPayments from 'dodopayments';
import OpenAI from 'openai';

const client = new DodoPayments({
  bearerToken: process.env.DODO_PAYMENTS_API_KEY,
});
const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });

export async function handleApiRequest(req, res) {
  const { customerId, prompt, model } = req.body;

  try {
    // 1. Check credit balance
    const balance = await client.creditEntitlements.balances.retrieve(customerId, {
      credit_entitlement_id: 'cde_your_credit_entitlement_id',
    });

    if (Number(balance.balance) <= 0) {
      return res.status(402).json({ error: 'Insufficient credits. Please top up.' });
    }

    // 2. Call OpenAI
    const completion = await openai.chat.completions.create({
      model: model,
      messages: [{ role: 'user', content: prompt }],
    });

    const { prompt_tokens, completion_tokens } = completion.usage;

    // 3. Ingest usage events to Dodo
    await client.usageEvents.ingest({
      events: [
        {
          event_id: `req_${completion.id}_in`,
          customer_id: customerId,
          event_name: 'llm.input_tokens',
          timestamp: new Date().toISOString(),
          metadata: { model, tokens: prompt_tokens }
        },
        {
          event_id: `req_${completion.id}_out`,
          customer_id: customerId,
          event_name: 'llm.output_tokens',
          timestamp: new Date().toISOString(),
          metadata: { model, tokens: completion_tokens }
        }
      ]
    });

    // 4. Return response to user
    res.json(completion);

  } catch (error) {
    console.error('API Error:', error);
    res.status(500).json({ error: 'Internal server error' });
  }
}
```

## Handling Edge Cases

A billing system like OpenAI's has several edge cases to plan for.

### Race Conditions

A user with a low balance can send several requests at once and exceed the balance before any event is processed. To prevent this, keep a small buffer, or hold a distributed lock on the customer's balance during each request.

### Event Ingestion Latency

Dodo Payments deducts credits asynchronously. A background worker processes new events about once a minute, so a deduction can lag the API call. For strict real-time enforcement, keep a local cache of each user's balance and update it as you serve requests.

### Refund Handling

Refunding a credit pack purchase doesn't remove the credits it granted. When you refund, deduct those credits yourself with a debit: use **Apply Credit/Debit** on the customer's **Credits** tab, or the [Create Ledger Entry](/api-reference/credit-entitlements/create-ledger-entry) API. Then update your application's view of the balance, so users can't spend credits they no longer have.

### Multi-Model Support

To support several models with different prices, choose one of two options:

1. **Separate Meters:** Create one set of meters per model, for example `gpt-4o.input_tokens` and `gpt-4o-mini.input_tokens`, each with its own **Meter units per credit**.
2. **Weighted Events:** Use one meter and multiply `tokens` by a weight before sending the event. For example, if GPT-4o costs 10 times as much as GPT-4o-mini, send 10 times the tokens for GPT-4o requests.

OpenAI publishes a separate rate for each model, and separate meters map to that structure most directly.

## Architecture Overview

The loop below shows the prepaid flow from purchase to blocked calls:

```mermaid theme={null}
flowchart LR
    A[Buy Credit Pack] --> B[Credits Added to Balance]
    B --> C[API Request]
    C --> D[Count Tokens]
    D --> E[Deduct Credits]
    E --> F{Balance > 0?}
    F -->|Yes| C
    F -->|No| G[Block API Calls]
    G --> A
```

The meters track tokens and deduct their dollar value from the user's credit balance at your configured rates. Your application blocks calls when the balance reaches zero.

## Conclusion

With Dodo Payments, you can combine usage-based billing with the predictability of prepaid credits, as OpenAI does. Customers pay up front, spend as they go, and top up when they need more.

The same pieces work for a large LLM platform or a small AI tool: a fiat credit, top-up products, token meters, and a balance check before each request.

## Key Dodo Features Used

These Dodo Payments features power the implementation:

<CardGroup cols={2}>
  <Card title="Credit-Based Billing" icon="coins" href="/features/credit-based-billing">
    Manage prepaid fiat credits and entitlements for your users.
  </Card>

  <Card title="Usage-Based Billing" icon="chart-line" href="/features/usage-based-billing/introduction">
    Track granular usage like tokens and bill for it.
  </Card>

  <Card title="One-Time Payments" icon="credit-card" href="/features/one-time-payment-products">
    Sell credit packs and top-ups through checkout.
  </Card>

  <Card title="Event Ingestion" icon="bolt" href="/features/usage-based-billing/event-ingestion">
    Send high-volume usage data to Dodo Payments.
  </Card>

  <Card title="Webhooks" icon="webhook" href="/developer-resources/webhooks/intents/credit">
    Stay updated on credit balance changes and low balance alerts.
  </Card>

  <Card title="LLM Ingestion Blueprint" icon="brain-circuit" href="/developer-resources/ingestion-blueprints/llm">
    Automatic token tracking for OpenAI and other LLM providers.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.