Blog · Field Notes
How to Benchmark LLM API Performance: TTFT, Streaming Stability, and a Reproducible Method (2026)
High performance is the most claimed and least verified phrase in the LLM API market. Three metrics actually measure it - TTFT P95 distribution, stream truncation rate, and cache-hit amplification - and here is a test you can run yourself, starting from one curl command.
12 min read→Enterprise AI API Security Checklist (2026): Key Governance, Zero Content Retention, Audit Trails, Vendor Verification
Evaluating the security of an LLM API vendor comes down to four layers: key governance, the data path, audit and reconciliation, and vendor verifiability. Here is a checklist where every item includes how to actually verify it - usable against any vendor, including us.
13 min read→How an LLM API Gateway Survives High Concurrency: Multi-Upstream Scheduling, Failover, and Cache Affinity (2026)
High concurrency for LLM APIs is a scheduling problem, not a throughput problem: upstream quotas are scarce, upstream errors are routine, and every request is a long-lived stream. Here is how an enterprise gateway actually handles it, plus a reproducible checklist to verify any vendor.
14 min read→How to Use DeepSeek Harness (dsh) with One Gateway Key: DeepSeek, GLM, Grok and Claude (2026)
DeepSeek's official agent runtime dsh passed 100k GitHub stars in two days. A three-minute HopBase setup: settings.yaml, protocol cheat sheet, verified function calling and streaming, quick troubleshooting.
6 min read→How to Run Claude Code for a Team: Access Paths and a Verifiable Gateway Checklist (2026)
Getting Claude Code / Codex into production as a team: three access paths compared, plus a 7-point gateway checklist you can verify yourself — protocol fidelity, model authenticity, failover, billing transparency, per-member keys.
8 min read→How to Use Oumomo: Six AI Tools from Viral Remake to Official TikTok Publishing (with the FastMoss Workflow)
A hands-on guide to Oumomo's six AI tools — viral video remake, URL to video, TikTok listing images, official API publishing, script generator and Amazon SEO — plus the FastMoss research workflow and partner discount.
8 min read→MiniMax Hailuo H3 Is Live on HopBase: Text, Image & Reference to Video with Native Audio, 15% Off List
MiniMax's new multimodal video model H3 is live: 4-15s clips at 768P/2K with native stereo audio and TTS across 11 languages, at 85% of the official list price.
5 min read→GLM 5.3 Is Live on HopBase: 1M Context, Thinking-Length Control, ~15% Off List
Zhipu's latest flagship GLM 5.3 is now available on HopBase via the Tencent official channel: 1M context, thinking-length control, base price aligned with the official list and a limited-time ~15% discount. Specs, pricing, and a three-minute quickstart inside.
4 min read→Kling Digital Human API: Turn a Photo and an Audio Clip into a Talking-Head Video
Two pipelines explained: kling-avatar generates a talking-head video from photos plus audio; kling-lip-sync re-voices existing footage. Full curl walkthrough, billing rules, and 25% off list.
11 min read→How to Verify Your LLM API Provider Isn't Swapping Models: A Reproducible Fidelity Test
Asking a model "who are you" proves nothing. Verify any API channel with cost-structure checks, model-family fingerprints, and blind quality benchmarks — including ours.
11 min read→Seedance API Pricing Explained (2026): Cost per Second, 2.0 vs 2.5, and How to Pay Less
Seedance bills by video token, not per clip. Here is every tier's rate, what a 15-second clip actually costs, and which tier to pick for testing, ad variants, and delivery.
7 min read→Tencent WorkBuddy Enterprise Pricing (2026): Plans, Discounts, and Onboarding
Flagship at ¥198 and Dedicated at ¥316 per user per month, both with 2,000 monthly Credits — and how partner-channel pricing takes 20% off the identical official product.
8 min read→Claude Code Setup in China: Complete Guide — from Registration to First Autocompletion in 10 Minutes
No VPN needed, no code changes — just two environment variables to connect Claude Code to a stable gateway. This guide covers registration, getting your API key, configuration, and verification, plus the three most common pitfalls.
7 min read→GPT-5.6 vs Claude Sonnet 5 for Chinese Project Code: Which Should You Choose?
We route both flagship coding models at scale through our production gateway. Here's what we actually observe about completion quality, instruction following, and token consumption — plus our real-world selection advice.
6 min read→How Much Can Prompt Cache Save You? Mechanics, Real-World Hit Rates, and Three Writing Techniques
Cache hit rate determines your bill. Here's how prompt caching actually works, the real hit rates we see in conversational tools, and three prompt organization techniques that can double your cache efficiency.
6 min read→Drop PDFs and Excel Files Straight Into Your Model: The Right Way to Use Attachments in AI Chat
Spreadsheets, scanned PDFs, raw email, web links — which ones go straight to the model, and which need preprocessing? Here's everything about AI Chat's attachment capabilities and how to use them effectively.
6 min read→Seedance 2.0 Video Generation API: Quick Start Guide
From submitting your job to downloading the final video: request parameters, task polling, video retrieval, and how resolution, duration, and pricing line up.
5 min read→Routing Multiple Upstreams with One API Key: Group-Based Routing Can Cut Costs in Half
Choose your tier per task: stable group during peak hours, economy group for batch jobs. Here's how groups, multipliers, and automatic failover work together, plus cost-optimized setups for three typical user profiles.
6 min read→










