SoraClip
2026.09.29
Assessment report 02

SoraClip Usage Management and Cost Estimates

Updated: September 29, 2026.

Currency: US dollars. Taxes, exchange-rate differences, and commercial discounts are excluded.

1. Purpose

Establish usage and cost statistics for SoraClip translation and the Artificial Intelligence Agent, referred to below as Agent, to support operating budgets, service-scale assessment, and future subscription plans.

This report covers required services, usage recording methods, and monthly cost estimates, using costs per user, per feature, and for the whole service as management indicators.

2. Services used

Service Purpose Cost basis
Soniox, real-time model stt-rt-v5 Speech recognition and text translation for the translation feature; real-time speech recognition for Agent Input audio, input text, and output text usage
OpenAI Realtime, model gpt-realtime-2.1 Generate Agent text and spoken responses from finalized Soniox transcripts Text input, text output, and audio output usage
Managed backend service Verify login identity, issue temporary credentials, receive usage, and aggregate costs Service runtime, data transfer, and related platform fees
Managed PostgreSQL database Centrally store user usage, rates, and cost records Database compute, storage, and backups
Scheduled jobs Regularly reconcile provider usage and update cost statistics Job execution fees

PostgreSQL is a relational database for storing structured data. The backend provides services through an Application Programming Interface (API) and uses sideband, a separate control connection, to receive OpenAI session events and usage. OpenAI documentation: Sideband.

The app connects directly to Soniox for recognition and translation. Agent then sends finalized transcripts to OpenAI, whose responses return directly to the app. The backend manages production service keys, while the app uses temporary credentials. Usage is attributed to the logged-in user. Platform configurations and fees are listed in the Backend Platform Comparison.

3. Recording methods

3.1 Record sources and timing

Source Recording method Update timing
Soniox Retrieve provider usage records and attribute them using previously associated user and feature information Query after use ends, with continued reconciliation through scheduled jobs
OpenAI Realtime Receive usage for each model generation through backend sideband, without waiting for the entire session to end After usage is finalized for each model generation
Backend platform Aggregate platform runtime, database, and other service fees Update from platform usage and monthly bills

Soniox provides model, audio duration, usage, and cost records for each request; OpenAI provides categorized usage for each model generation. Cost statistics use data returned by the providers, without counting the same usage twice. Soniox: Usage Logs, OpenAI: Realtime usage.

3.2 Stored information

Information Contents Management use
User Internal account identifier Calculate individual and average usage costs
Feature Translation or Agent Analyze usage and cost share by feature
Date Usage date and month Produce daily and monthly statistics and trends
Raw usage Audio duration, text and audio usage, and cached usage Preserve the basis for cost calculations
Rate Provider, model, unit price, currency, and effective date Calculate and compare costs using the correct rates
Cost Model fees, platform fees, and reconciliation status Compare budgeted and actual spending

Raw usage and calculated amounts are stored separately, preserving the basis for historical calculations when rates change. Unreconciled amounts remain marked as pending; statistics are updated after confirmation.

3.3 Aggregation

Model fees are summarized by user, feature, and month, with backend platform fees listed separately to provide these management indicators:

  • Total monthly costs and budget variance.
  • Cost shares for translation and Agent.
  • Average cost per active user.
  • Cost distribution among users with high usage.

User costs reflect their model consumption. Shared platform costs are listed separately and allocated by active user count or usage when needed, keeping them distinct from direct usage costs.

4. Cost estimates

4.1 Calculation basis

This section uses public usage-based rates checked on September 29, 2026. A token is a provider's unit for measuring model usage, not a word count. Token counts from different providers and audio models are not directly interchangeable.

Provider / model Billing item Price per million tokens
Soniox stt-rt-v5 Input audio $2.00
Soniox stt-rt-v5 Input text, output text $4.00 each
OpenAI gpt-realtime-2.1 Text input, uncached $4.00
OpenAI gpt-realtime-2.1 Text output $24.00
OpenAI gpt-realtime-2.1 Audio output $64.00
OpenAI gpt-realtime-2.1 Cached text input $0.40

Rate sources: Soniox official pricing, OpenAI official model pricing. Caching means the provider reuses previously processed input, which qualifies for a lower rate when applicable. This example does not apply caching discounts.

4.2 Usage assumptions

This report and the platform comparison use the same usage assumptions per person for 100, 500, and 1,000 monthly active users. Frequency, duration, and peak load are estimates, not measurements, and will be adjusted using usage records after launch.

Usage per person Monthly assumption
Translation 30 sessions × 10 minutes each = 300 minutes
Agent conversations 10 conversations × 5 minutes each = 50 minutes
Agent responses 10 responses per conversation, 100 in total
User speech in Agent 1 minute per conversation, 10 minutes in total
Soniox audio streaming for Agent Full 5 minutes per conversation, 50 minutes in total
sideband 1 connection throughout each Agent conversation, 50 minutes in total
Monthly usage / peak load 100 users 500 users 1,000 users
Translation sessions 3,000 15,000 30,000
Translation hours 500 hours 2,500 hours 5,000 hours
Agent conversations 1,000 5,000 10,000
Agent responses 10,000 50,000 100,000
User speech hours in Agent 16.67 hours 83.33 hours 166.67 hours
Soniox streaming / sideband connection hours for Agent 83.33 hours each 416.67 hours each 833.33 hours each
Peak sideband connections 10 50 100
App authorization requests 4,000 20,000 40,000
Usage records 14,000 70,000 140,000

Usage records follow the session and response units above; individual streaming chunks are not stored. Connection hours sum the duration of all sideband connections, rather than the runtime of a single service.

Soniox input audio is estimated at 30,000 tokens per hour. Translation assumes 15,000 text output tokens per hour each for recognition and translation. Agent recognition text output is based on actual speaking time at 15,000 tokens per hour. Agent audio input conservatively assumes continuous transmission throughout the conversation; waiting time is not treated as free. These audio and recognition conversions reference official figures; translation text length is an assumption for this example. Soniox: Usage conversion.

OpenAI receives finalized Soniox text as the prompt. This estimate uses text input and text conversation history, with no audio input fee. Monthly usage per user assumes 100,000 uncached text input tokens, 10,000 text output tokens, and 10,000 audio output tokens. Text input already includes system instructions and text conversation history, so they are not added again. This is an assumption for the budget configuration; if actual sessions retain audio history, that usage must be added according to provider settlement records. OpenAI: Usage and multi-turn costs.

Response counts describe usage scale. Fees still depend on actual usage from prompts, responses, conversation history, and tools.

4.3 Monthly cost calculation

Monthly fee per item = monthly usage per person × active users ÷ 1,000,000 × applicable unit price.

Item Monthly tokens per person 100 users 500 users 1,000 users
Soniox translation: input audio 150,000 $30.00 $150.00 $300.00
Soniox translation: recognition and translation text output 150,000 $60.00 $300.00 $600.00
Soniox Agent recognition: input audio 25,000 $5.00 $25.00 $50.00
Soniox Agent recognition: text output 2,500 $1.00 $5.00 $10.00
OpenAI Agent: text input 100,000 $40.00 $200.00 $400.00
OpenAI Agent: text output 10,000 $24.00 $120.00 $240.00
OpenAI Agent: audio output 10,000 $64.00 $320.00 $640.00
Total model fees — $224.00 $1,120.00 $2,240.00

Translation costs approximately $0.18 per hour in this example, matching the reference figure on the official translation page. Actual billing still depends on usage such as output text, rather than a fixed hourly package. Soniox: Real-time translation pricing.

The average model cost per active user is $2.24/month, comprising $0.90 for translation and $1.34 for Agent (Soniox recognition $0.06 and OpenAI responses $1.28). All three scenarios use the same usage per user, so model fees increase in proportion to the number of users.

4.4 Combined model and platform costs

Combined monthly platform and model cost = Soniox + OpenAI model fees + the selected backend platform fee. Additional tools, operational labor, and applicable taxes are charged separately.

Using the initial configurations in the Platform Comparison, the combined monthly model and backend platform fees are:

Platform and cost 100 users 500 users 1,000 users
Soniox + OpenAI model fees $224.00 $1,120.00 $2,240.00
Render backend platform fee $73.00 $73.00 $76.75
Render platform + model total $297.00 $1,193.00 $2,316.75
Render average per user/month $2.97 $2.39 $2.32
Microsoft Azure backend platform fee $195.68 $195.68 $210.63
Microsoft Azure platform + model total $419.68 $1,315.68 $2,450.63
Microsoft Azure average per user/month $4.20 $2.63 $2.45

Platform fees combine fixed configurations with traffic and log allowances for each scale: outbound transfer of 5, 25, and 50 GB/month, and logs of 1, 5, and 10 GB/month. These allowances are not measured usage. Both Render and Azure deduct included or free allowances assumed to remain available. Fixed compute and database fees do not rise in direct proportion to user count, so the platform cost allocated per user decreases at larger scales.

Platform estimates assume two company administrators and one application and database in a single region, without high-availability redundancy. Azure uses a General Purpose database with more compute and storage than the Render configuration in this example; specifications and pricing conditions are in the platform comparison. These totals exclude search and other tools, additional Soniox text context input, taxes, development and maintenance labor, and payment processing fees.

More users, longer responses, more conversation history, and longer billable streaming time all increase costs. Cache hits may reduce OpenAI input costs. After launch, assumptions should be updated with actual usage and budget variance reviewed monthly. This example is not a cost ceiling.

4.5 Average cost per user

Average monthly cost per user = combined platform and model fees ÷ monthly active users. For example, at 1,000 users, Render is $2,316.75 ÷ 1,000 = $2.32/user/month, and Microsoft Azure is $2,450.626 ÷ 1,000 = $2.45/user/month, both rounded to two decimal places.

Cost per user/month 100 users 500 users 1,000 users
Translation model fees $0.90 $0.90 $0.90
Agent model fees $1.34 $1.34 $1.34
Allocated Render platform cost $0.730 $0.146 $0.077
Allocated Microsoft Azure platform cost $1.957 $0.391 $0.211

The two platforms are alternatives; their allocated fees are not added together. The denominator is users active during the month. Allocated platform costs use three decimal places, and other amounts use two. Totals are calculated from unrounded values.

5. Future subscription integration

For now, record usage and costs with user quotas set to "unlimited" where appropriate. Future subscription plans will define the measurement unit, quota per period, validity period, and handling of excess usage. The same usage records will then update remaining quotas.

Management item Current phase After subscriptions launch
Usage and cost records Continue recording and aggregating by user and feature Use the same records
Usage quota Explicitly mark as unlimited Set quotas and periods by subscription plan
Remaining quota Do not interpret unlimited as zero quota Update from the plan limit and usage to date
Cost analysis Assess average costs and the distribution of high usage Assess plan costs, gross margins, and quota settings

Provider costs and user quotas are managed separately, so subscription metering is not constrained by a single provider's billing unit.