The bill scales with the workload
Invoice parsing, field extraction, classification, summarization — high volume, low reasoning. These are exactly the jobs where the model tier stops mattering and only the price does.
Swap one line — the base_url — and run the same workloads on
DeepSeek V4.1 Flash. We handle the contract, the invoice, and the account contact.
You keep the savings.
from openai import OpenAI client = OpenAI(- base_url="https://api.deepseek.com/v1",- api_key=os.environ["DEEPSEEK_API_KEY"],+ base_url="https://api.workhorseapi.com/v1",+ api_key=os.environ["WORKHORSE_API_KEY"],) # everything else stays the same
Built for teams running invoice parsing, document extraction, and classification at scale.
It usually starts the same way: someone picks the strongest model available, ships it, and it works. Two years later the workload has grown tenfold — and nobody has gone back to ask whether it still needs the strongest model.
Invoice parsing, field extraction, classification, summarization — high volume, low reasoning. These are exactly the jobs where the model tier stops mattering and only the price does.
Self-service APIs are built for individuals: no account manager, a standard agreement you can't amend, prepaid top-up instead of a purchase order, a sign-up that rejects corporate email domains.
Engineering has a roadmap. Re-plumbing a working pipeline to save money is nobody's priority — until someone is asked why the AI line item doubled.
Everything below comes from the model's published specification — no benchmark theater, no invented numbers.
Feed it a whole contract, a full claims file, or a multi-hundred-page PDF without building a chunking pipeline first.
1,000,000 tokensStructured extraction at scale — long JSON arrays, full document rewrites, bulk classification in a single response.
384,000 tokensReceipts, scans, screenshots, and footage go in directly — no separate OCR stack to run and pay for alongside the model.
vision inputBatch jobs, nightly backfills, and re-processing runs move to off-peak hours and cost half the peak rate. Published pricing, not a promotion.
50% off-peakSCANNED INPUT
read natively
Sample layout. No real invoice data.
STRUCTURED OUTPUT
image in, fields out
{
"vendor": "…",
"invoice_number": "…",
"issue_date": "…",
"line_items": [
{ "description": "…", "qty": … },
{ … }
],
"tax": …,
"total": …
}
Field names shown; values omitted.
No migration project. No procurement odyssey. No change to your application beyond the endpoint.
Send us your monthly token usage and what you're running today. If we can't beat your current cost, we'll say so.
A named account contact, a contract your legal team can negotiate, and a proper invoice for your finance team.
Point your existing OpenAI-compatible client at our endpoint. Keep your current provider configured as a fallback.
WE HANDLEYOU DO
1
Tell us your volume
Monthly token usage, and what you are running today. If we cannot beat your current cost, we say so.
2
We quote and contract
A named account contact, a negotiable agreement, an invoice.
- base_url="…"
+ base_url="…"
# nothing else changes3
You change one line
Point your existing client at our endpoint. Keep a fallback.
base_url.Diagram of the onboarding sequence. Not a screenshot.Each of these has its own page with the full setup — the workload shape, the API call, and what it costs against what you run today.
Line items, tax, totals, and vendor details pulled out of supplier invoices — including scanned and photographed ones.
Read the invoice parsing pageTurn contracts, claims files, and forms into structured records — long documents go in whole, no chunking pipeline required.
Read the document extraction pageSort inbound tickets, emails, and documents into the right queue — high volume, and the accuracy bar sits well within Flash territory.
Read the classification pagePhotographed receipts and low-quality scans read directly by the model — no separate OCR service to run alongside it.
Read the receipt reading pageCompletion, review, and test generation inside your own tooling — one endpoint, so it drops into what your team already uses.
Read the code assistance pageDatasets generated from your own seed documents and rules — not invented from a bare prompt, and validated against your schema.
Read the synthetic data pageThese are the six workloads we see most. If yours isn't here, describe it in the form below — or read how we compare to going direct.
The person who wants it and the person who can stop it are usually different people. Here is what each of them gets to hear.
“Can we amend the terms?”
We're not limited to one fixed standard agreement. We sign a data processing agreement, and we'll work through your legal team's questions before you commit.
“How does this get booked?”
Bank transfer or wire. One monthly invoice with the detail your accounting team needs. Net terms available for established accounts.
“What if it doesn't work out?”
We're a second channel, not a replacement. Leave your existing configuration in place and route a share of traffic to us. You can move it back the same way you moved it over.
ONE LINE CHANGES
Your application
The OpenAI SDK you already ship.
api.workhorseapi.com
Same request shape. Same streaming.
WORKHORSE ENDPOINT
DeepSeek V4.1 Flash
1M context · 384K output
response returns the same way
WHAT WE HANDLE AROUND IT
That is the anchor for every quote. Where your own figure lands depends on volume, cache hit ratio, and how much of the load can move off-peak — tell us those and we'll send it.
Tell us your monthly token volume and what you're running today. We'll come back with a rate and a straight comparison against your current bill.
Get a quoteNo. We're an independent provider — not affiliated with or endorsed by DeepSeek. We resell access to DeepSeek V4.1 Flash and handle the commercial side around it: contracting, invoicing, settlement, and a named account contact.
If you're happy with self-service, going direct is a perfectly reasonable option.
By bank transfer or wire against a monthly invoice. There is no checkout on this site — we don't take card payments here. Pricing is agreed with your account contact and confirmed in the contract.
Yes, both. Contract terms are negotiable, and we can sign a data processing agreement. That's the main practical difference from a self-service API, where the standard agreement typically can't be amended.
Usually the same day. The API is OpenAI-compatible, so in most codebases it's a change to base_url and the API key — nothing else. We'll walk your engineers through it.
We process request data to serve your requests and don't sell it. Full detail is in our privacy policy and data processing agreement. Tell us if you have residency requirements and we'll confirm what we can support before you commit.
Off-peak calls are billed at half the peak rate. It suits batch workloads — overnight backfills, bulk re-processing, periodic extraction runs — where a delay of a few hours costs you nothing.
OpenRouter routes across a very wide range of models. We do one thing: DeepSeek V4.1 Flash, as cheaply as we can, with the commercial process handled. If you need a dozen model families behind one key, they're the better fit. If DeepSeek is doing the work, we're usually the cheaper route.
You get a named account contact — a person, not a ticket queue. Keep your current provider configured as a fallback so a bad day doesn't become an outage.
It's a Flash-class model, not a frontier reasoning model. Complex multi-step reasoning and hard math are better served by a larger model. It also doesn't take audio input. If your workload needs either, we'll tell you rather than sell you something that won't fit.
Yes — the model handles multilingual workloads well, including translation and extraction from non-English documents. Tell us which languages matter to you and we'll confirm before you commit.
Yes. The endpoints follow the OpenAI chat-completions schema, so the official openai SDKs — and anything built on them — work unchanged. For most teams the whole migration is base_url plus the API key.
V4.1 Flash takes a 1M-token context window and returns up to 384K tokens. That covers long documents, multi-file jobs, and large extraction batches without hand-chunking everything.
Concurrency on Flash goes up to 2,500. We set your account's limits when we onboard you, based on the volume you tell us — give us your peak concurrency and we'll confirm in writing what we can commit to.
Yes, and batch is where the economics get interesting. Pair it with off-peak and the per-token rate is half the peak rate. It suits overnight backfills and periodic re-processing, where a few hours of delay costs you nothing.
Cached input tokens are billed at a far lower rate than fresh ones. If you re-send the same large prefix on every call — a schema, a long instruction block, a fixed set of examples — caching is usually the single biggest lever on the bill. We'll look at your prompt structure with you.
That's what most of our customers run on it. Flash takes image input natively, so scans of invoices, receipts, and forms go in directly — no separate OCR step in front of it. Video input works too.
Yes to both. We can sign a data processing agreement, and the master agreement is negotiable. A self-service API typically can't amend its standard terms — that's usually the point where a legal review stops.
We'll put response and availability commitments in the contract, and tell you plainly what we can and can't guarantee before you sign. We don't publish a blanket uptime number that we'd then have to defend.
You get a named contact and a monthly itemized invoice. Tell us what your finance team needs to see on it and we'll shape the reporting around that.
We screen companies and end users against applicable export controls and sanctions before onboarding. Tell us where your entity and your users are and we'll confirm before you commit to anything.
Nothing holds you. Your integration is a base_url and a key — the same code runs against any OpenAI-compatible endpoint. There's no proprietary SDK to unwind.
Tell us your monthly volume and what you're running today, using the form below. We'll usually come back within one business day with a price and a named contact.
One of us reads every submission and usually replies within one business day — with a rate and a comparison against your current bill.