性能
LW AI is designed for high throughput and low latency. This page provides performance benchmarks and optimization guidance.
Latency Benchmarks
Typical response times measured from the Hillsboro and Singapore regions:
| Model | Time to First Token (TTFT) | Tokens/sec (Output) |
|---|---|---|
| deepseek/deepseek-v4-flash | ~200ms | ~80 tok/s |
| deepseek/deepseek-v4-pro | ~500ms | ~40 tok/s |
| z-ai/glm-5.1 | ~400ms | ~50 tok/s |
| z-ai/glm-5.2 | ~600ms | ~35 tok/s |
| linkwo/fusion | ~300ms | ~60 tok/s |
| astra | ~400ms | ~45 tok/s |
Values are approximate and vary based on input length, output length, and current load. Measured with short prompts (~100 tokens input).
Regional Latency
| Client Location → API Region | Estimated Latency |
|---|---|
| Singapore → openapi.linkwo.ai | ~15ms |
| Tokyo → openapi.linkwo.ai | ~50ms |
| US West → openapi.linkwo.ai | ~15ms |
| US East → openapi.linkwo.ai | ~60ms |
| Europe → openapi.linkwo.ai | ~200ms |
| South America → openapi.linkwo.ai | ~120ms |
Throughput
| Tier | Max Concurrent Requests | Estimated Max TPM |
|---|---|---|
| Free | 2 | 40,000 |
| Standard | 10 | 200,000 |
| Pro | 30 | 1,000,000 |
| Enterprise | Custom | Custom |
Optimization Tips
Reduce Latency
- Choose the right region — Use
openapi.linkwo.ai(routes to the nearest region automatically) - Use Flash models — DeepSeek V4 Flash is 2-3x faster than Pro
- Enable streaming — Receive tokens as they're generated instead of waiting for the full response
- Shorten prompts — Less input tokens means faster processing
Maximize Throughput
- Batch requests — Use the Batch API for non-real-time workloads
- Connection pooling — Reuse HTTP connections across requests
- Concurrent requests — Send multiple requests in parallel (within rate limits)
- Cache results — Avoid duplicate API calls for identical inputs
Reduce Token Usage
- Set
max_tokens— Limit output length when you don't need long responses - Use concise prompts — Be specific and avoid unnecessary context
- Leverage system prompts — Set behavior once instead of repeating in every message
- Use structured output — Request JSON directly to avoid verbose formatting
Monitoring Your Usage
Contact support@linkwo.com to get performance metrics for your account:
- Request latency percentiles (p50, p90, p99)
- Token usage over time
- Error rates by model
- Rate limit utilization
Related
- Rate Limits — API rate limiting
- Routing & Fallback — Automatic fallback
- Service Status — Real-time status
