> ## Documentation Index
> Fetch the complete documentation index at: https://docs.dodopayments.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Replicate Billing Model

> Deconstruct Replicate's per-second hardware-specific billing and build the same pure usage-based model using Dodo Payments.

Replicate is a platform for running open-source machine learning models in the cloud, and one of the purest examples of usage-based pricing in AI. There is no monthly subscription and no flat rate per model run. Most models are billed for the compute time they use, measured per second, at a rate that depends on the hardware. Some models are billed by input and output instead, such as per image or per token ([Replicate pricing](https://replicate.com/pricing)).

Per-second billing suits AI workloads because run times are unpredictable. One user might run a lightweight model for a few seconds, and another a large generative model for several minutes. Because the price follows the compute used rather than the model, it stays transparent at any scale.

## How Replicate Bills

Replicate's time-based pricing doesn't depend on the model. Whether you generate an image with SDXL or run Llama 3, the bill depends on the hardware tier and the run time. Replicate can host thousands of open-source models without a separate price for each.

| Hardware | Price per Second | Price per Hour |
| :- | :- | :- |
| CPU | \$0.000100 | \$0.36 |
| NVIDIA T4 GPU | \$0.000225 | \$0.81 |
| NVIDIA L40S GPU | \$0.000975 | \$3.51 |
| NVIDIA A100 (80GB) GPU | \$0.001400 | \$5.04 |
| NVIDIA H100 GPU | \$0.001525 | \$5.49 |

The implementation below also uses example meters for A40 and A100 (40GB) tiers, which Replicate's pricing page no longer lists. Create one meter for each hardware tier you offer.

Each run is metered on the meter for the hardware it used:

```mermaid theme={null}
flowchart TD
    A[Model Execution Completes] --> B[Measure Execution Time]
    B --> C{Which Hardware?}
    C -->|CPU| D[CPU Meter]
    C -->|T4 GPU| E[T4 GPU Meter]
    C -->|A40 GPU| F[A40 GPU Meter]
    C -->|A40 Large GPU| G[A40 Large GPU Meter]
    C -->|A100 40GB GPU| H[A100 40GB GPU Meter]
    C -->|A100 80GB GPU| I[A100 80GB GPU Meter]
    D --> J[Bill per Second at Tier Rate]
    E --> J
    F --> J
    G --> J
    H --> J
    I --> J
```

1. **Hardware-Specific Rates**: The price per second depends on the compute resources used. Each hardware tier has its own rate.
2. **Pure Usage-Based Model**: There are no monthly fees, no included quotas, and no overages. Users pay for exact compute time, for example "12.4 seconds on an A100", not per generation.
3. **Per-Second Granularity**: Providers that bill by the hour or minute charge for unused time on short tasks. Per-second billing removes that waste for small experiments and large production workloads alike.

<Info>
  For public models, Replicate charges only for the time a model is actively processing your request. Setup (cold boot) and idle time are free. Private models and deployments run on dedicated hardware, so you pay for all the time their instances are online: setup, idle, and active ([Replicate billing docs](https://replicate.com/docs/topics/billing)).
</Info>

## What Makes It Unique

* **Hardware-specific metering:** The same model costs more on faster hardware, so users trade speed against cost. A T4 GPU suits tasks that aren't time-sensitive, and an A100 suits real-time applications.
* **Per-second granularity:** Billing is calculated to the second, so users don't pay for unused time on short tasks.
* **No subscription:** Users start with no commitment, and cost scales with usage. This suits startups and developers trying different models.
* **Model-agnostic:** The same billing logic applies to image generation, text processing, audio transcription, and video synthesis. The platform supports a large model catalog without complex pricing tables.

## Build This with Dodo Payments

You can build this model with Dodo Payments usage-based billing. Create one meter per hardware tier and attach all of them to a single product.

<Steps>
  <Step title="Create Usage Meters (One Per Hardware Class)">
    Create a separate meter for each hardware tier. Each tier has its own price per second, so separate meters let Dodo Payments price each tier and itemize the invoice.

    | Meter Name | Event Name | Aggregation | Property |
    | :- | :- | :- | :- |
    | CPU Compute | `compute.cpu` | Sum | `execution_seconds` |
    | GPU T4 Compute | `compute.gpu_t4` | Sum | `execution_seconds` |
    | GPU A40 Compute | `compute.gpu_a40` | Sum | `execution_seconds` |
    | GPU A40 Large Compute | `compute.gpu_a40_large` | Sum | `execution_seconds` |
    | GPU A100 40GB Compute | `compute.gpu_a100_40` | Sum | `execution_seconds` |
    | GPU A100 80GB Compute | `compute.gpu_a100_80` | Sum | `execution_seconds` |

    **Sum** aggregation over `execution_seconds` gives the total compute time per hardware tier for the billing period.
  </Step>

  <Step title="Create a Usage-Based Product">
    Create a product in the Dodo Payments dashboard with these settings:

    * **Pricing type:** Usage Based Billing
    * **Base Price:** \$0/month (no subscription fee)
    * **Billing frequency:** Monthly

    Add every meter with its price per unit:

    | Meter | Price Per Unit (per second) |
    | :- | :- |
    | compute.cpu | \$0.000100 |
    | compute.gpu\_t4 | \$0.000225 |
    | compute.gpu\_a40 | \$0.000575 |
    | compute.gpu\_a40\_large | \$0.000725 |
    | compute.gpu\_a100\_40 | \$0.001150 |
    | compute.gpu\_a100\_80 | \$0.001400 |

    Set the **Free Threshold** to 0 on every meter, so every second of execution is billable.
  </Step>

  <Step title="Send Usage Events">
    Send a usage event to Dodo Payments when each model run completes. Give each prediction a unique `event_id`, so a retried event isn't counted twice.

    ```typescript theme={null}
    import DodoPayments from 'dodopayments';

    type HardwareTier = 'cpu' | 'gpu_t4' | 'gpu_a40' | 'gpu_a40_large' | 'gpu_a100_40' | 'gpu_a100_80';

    const client = new DodoPayments({
      bearerToken: process.env.DODO_PAYMENTS_API_KEY,
    });

    async function trackModelExecution(
      customerId: string,
      modelId: string,
      hardware: HardwareTier,
      executionSeconds: number,
      predictionId: string
    ) {
      const eventName = `compute.${hardware}`;

      await client.usageEvents.ingest({
        events: [{
          event_id: `pred_${predictionId}`,
          customer_id: customerId,
          event_name: eventName,
          timestamp: new Date().toISOString(),
          metadata: {
            execution_seconds: executionSeconds,
            model_id: modelId,
            hardware: hardware
          }
        }]
      });
    }

    // Example: SDXL image generation on A100
    await trackModelExecution(
      'cus_abc123',
      'stability-ai/sdxl',
      'gpu_a100_80',
      8.3,  // 8.3 seconds of A100 time
      'pred_xyz789'
    );
    ```
  </Step>

  <Step title="Measure Execution Time Precisely">
    Time each model run with `performance.now()`, and round to the nearest tenth of a second for billing.

    ```typescript theme={null}
    // executeModel is your application's own function that runs the model.
    async function runModelWithMetering(
      customerId: string,
      modelId: string,
      hardware: HardwareTier,
      input: Record<string, unknown>
    ) {
      const predictionId = `pred_${Date.now()}`;
      const startTime = performance.now();

      try {
        const result = await executeModel(modelId, input, hardware);
        const executionSeconds = (performance.now() - startTime) / 1000;
        const billedSeconds = Math.round(executionSeconds * 10) / 10;

        await trackModelExecution(
          customerId,
          modelId,
          hardware,
          billedSeconds,
          predictionId
        );

        return result;
      } catch (error) {
        // This example bills compute time for failed runs longer than 1 second.
        // Replicate doesn't charge for failed runs on public models; remove this block to match.
        const executionSeconds = (performance.now() - startTime) / 1000;
        if (executionSeconds > 1) {
          await trackModelExecution(
            customerId,
            modelId,
            hardware,
            Math.round(executionSeconds * 10) / 10,
            predictionId
          );
        }
        throw error;
      }
    }
    ```
  </Step>

  <Step title="Create Checkout">
    When a user signs up, create a checkout session for the usage-based product. Dodo Payments then bills the usage each cycle and issues the invoices.

    ```typescript theme={null}
    const session = await client.checkoutSessions.create({
      product_cart: [
        { product_id: 'pdt_compute_payg', quantity: 1 }
      ],
      customer: { email: 'ml-engineer@example.com' },
      return_url: 'https://yourplatform.com/dashboard'
    });
    ```
  </Step>
</Steps>

## Accelerate with the Time Range Ingestion Blueprint

The [Time Range Ingestion Blueprint](/developer-resources/ingestion-blueprints/time-range) shortens per-second compute tracking. Create one ingestion instance per hardware tier, and call `trackTimeRange` after each run.

```bash theme={null}
npm install @dodopayments/ingestion-blueprints
```

```typescript theme={null}
import { Ingestion, trackTimeRange } from '@dodopayments/ingestion-blueprints';

// Create one ingestion instance per hardware tier
function createHardwareIngestion(hardware: string) {
  return new Ingestion({
    apiKey: process.env.DODO_PAYMENTS_API_KEY,
    environment: 'live_mode',
    eventName: `compute.${hardware}`,
  });
}

const ingestions: Record<string, Ingestion> = {
  cpu: createHardwareIngestion('cpu'),
  gpu_t4: createHardwareIngestion('gpu_t4'),
  gpu_a40: createHardwareIngestion('gpu_a40'),
  gpu_a40_large: createHardwareIngestion('gpu_a40_large'),
  gpu_a100_40: createHardwareIngestion('gpu_a100_40'),
  gpu_a100_80: createHardwareIngestion('gpu_a100_80'),
};

// executeModel, modelId, input, hardware, and customerId come from your application.
const startTime = performance.now();
const result = await executeModel(modelId, input, hardware);
const durationMs = performance.now() - startTime;

await trackTimeRange(ingestions[hardware], {
  customerId: customerId,
  durationMs: durationMs,
  metadata: {
    model_id: modelId,
    hardware: hardware,
  },
});
```

The blueprint builds and sends the event, with the duration as `durationMs` in the metadata. Set each meter's **Over Property** to `durationMs`, or send `durationSeconds` to keep per-second pricing. With one ingestion instance per hardware tier, this maps directly to Replicate's multi-tier metering.

<Tip>
  For long-running jobs, combine the Time Range Blueprint with interval-based heartbeat tracking, shown in [Advanced: Heartbeat Metering](#advanced-heartbeat-metering). See the [full blueprint documentation](/developer-resources/ingestion-blueprints/time-range) for more patterns.
</Tip>

## Cost Estimation for Users

Usage-based bills can be hard to predict, so show users a cost estimate before they run a model. Estimates prevent surprise bills and build trust.

### Example Cost Calculations

| Model | Hardware | Avg Time | Cost Per Run |
| :- | :- | :- | :- |
| SDXL (image) | A100 80GB | \~8 sec | \~\$0.0112 |
| Llama 3 (text) | A100 40GB | \~3 sec | \~\$0.0035 |
| Whisper (audio) | GPU T4 | \~15 sec | \~\$0.0034 |

### Building a Cost Calculator

This function multiplies the tier's per-second rate by the estimated run time:

```typescript theme={null}
function estimateCost(hardware: HardwareTier, estimatedSeconds: number): number {
  const rates: Record<HardwareTier, number> = {
    'cpu': 0.000100,
    'gpu_t4': 0.000225,
    'gpu_a40': 0.000575,
    'gpu_a40_large': 0.000725,
    'gpu_a100_40': 0.001150,
    'gpu_a100_80': 0.001400
  };

  return Number((rates[hardware] * estimatedSeconds).toFixed(4));
}

// Show the user before running: "This will cost approximately $0.0119"
const estimate = estimateCost('gpu_a100_80', 8.5);
```

## Enterprise: Reserved Capacity

For customers who need dedicated capacity and no cold boots, Replicate offers deployments: dedicated instances billed for all the time they're online.

To model reserved capacity with Dodo Payments, sell it as a subscription product:

* **Product Type:** Subscription
* **Price:** Fixed monthly price (for example, "Reserved A100 Instance - \$500/month")
* **Billing Cycle:** Monthly

You can still send usage events for monitoring and analytics, while the subscription covers the cost. As a customer's volume grows, reserved capacity often costs less than pay-as-you-go.

## Advanced: Heartbeat Metering

For tasks that run for minutes or hours, one event at the end is risky: if the process crashes, you lose the usage data. Instead, send a usage event every 30 to 60 seconds while the task runs.

```typescript theme={null}
// executeLongTask is your application's own function. trackModelExecution is defined above.
async function runLongTaskWithHeartbeat(
  customerId: string,
  modelId: string,
  hardware: HardwareTier
) {
  const predictionId = `pred_${Date.now()}`;
  let totalSeconds = 0;

  const heartbeatInterval = setInterval(async () => {
    try {
      await trackModelExecution(
        customerId,
        modelId,
        hardware,
        30,
        `${predictionId}_${totalSeconds}`
      );
      totalSeconds += 30;
    } catch (error) {
      console.error('Heartbeat tracking failed:', error, { predictionId, totalSeconds });
    }
  }, 30000);

  try {
    await executeLongTask();
  } finally {
    clearInterval(heartbeatInterval);
  }
}
```

## Key Dodo Features Used

<CardGroup cols={2}>
  <Card title="Usage-Based Billing" icon="chart-line" href="/features/usage-based-billing/introduction">
    Set up products that bill based on consumption.
  </Card>

  <Card title="Meters" icon="gauge" href="/features/usage-based-billing/meters">
    Define the metrics you want to track and bill for.
  </Card>

  <Card title="Event Ingestion" icon="bolt" href="/features/usage-based-billing/event-ingestion">
    Send usage data to Dodo Payments as it happens.
  </Card>

  <Card title="Subscriptions" icon="calendar" href="/features/subscription">
    Manage recurring billing for reserved capacity and enterprise plans.
  </Card>

  <Card title="Time Range Blueprint" icon="clock" href="/developer-resources/ingestion-blueprints/time-range">
    Per-second compute tracking with duration helpers.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.