CrispASR

Concurrency, parallelism & scaling

How CrispASR uses multiple cores, how it handles concurrent requests, and how to scale to large batch workloads (thousands of hours of audio). If you skimmed the docs looking for “concurrency”, “parallel”, “batch”, or “throughput” and found nothing — this is the page.

TL;DR


The three layers of parallelism

CrispASR parallelises at three levels. They compose: a bulk run typically uses all three (many processes × server workers × intra-op threads only if you size them to your hardware — over-subscribing all three at once contends).

1. Intra-op threads (inside one transcription)

Every inference already runs multi-threaded inside ggml. One crispasr invocation transcribing one file uses several CPU cores for the matmuls/convs.

This is within a single stream — it does not let two files transcribe at once. On a GPU backend the heavy math runs on the device and thread count matters less.

2. The HTTP server: concurrent transport, serialized model

crispasr --server loads the model once and reuses it across requests (no per-request reload). The HTTP layer (cpp-httplib) has its own thread pool (≥8 threads), so it accepts and parses many requests concurrently.

But by default there is a single model instance, and inference is mutex-serialized. N simultaneous uploads are received in parallel, then run strictly one-at-a-time through the one model. This is deliberate: a single loaded model/context is not safe to use from multiple threads at once (see include/crispasr.h), so the server holds a model_mutex around the inference call.

So out of the box the server gives you a persistent, no-reload model with serialized execution — great latency per request, throughput of one stream.

3. --server-workers N: concurrent inference via independent instances

To run several transcriptions at the same time inside one server process, load more than one model instance:

# Flag form (shown in --help):
crispasr --server -m model.gguf --port 8080 --server-workers 4

# Env form (overrides the flag; the ultimate gate):
CRISPASR_SERVER_WORKERS=4 crispasr --server -m model.gguf --port 8080

Each worker owns its own backend instance (and, on CUDA, its own device context / command queue). A request is routed to a free worker via a blocking RAII lease from a bounded pool (src/core/worker_pool.h), so up to N run concurrently and different workers never contend.

Important caveats — read before setting this above 1:

Transcripts are identical whether a request ran on a pooled worker or the shared model — the pool changes scheduling, never the math.


Offline bulk transcription (the “thousands of hours” case)

If your goal is to get through a large corpus of files (not to serve live requests), you do not need the server at all. Bulk ASR is embarrassingly parallel across files, and the simplest robust approach is to run several crispasr processes at once, each transcribing whole files.

The CLI itself processes multiple input files sequentially (no -j flag), so drive the parallelism from the shell:

# N parallel processes, one file each. Tune -P to your hardware
# (see "sizing" below). Auto-download resolves the model once into the cache;
# point every process at the same cached model to avoid N downloads.
mkdir -p out
ls corpus/*.wav | xargs -P 4 -I{} \
  sh -c 'crispasr -m model.gguf --threads 4 -otxt -of "out/$(basename "$1" .wav)" -f "$1"' _ {}

# GNU parallel is nicer for progress, retries, and remote nodes
# (here {/.} = input basename without extension):
parallel -j 4 --bar \
  'crispasr -m model.gguf --threads 4 -otxt -of out/{/.} -f {}' ::: corpus/*.wav

Sizing -P × --threads. The two multiply. On a CPU box, total worker threads ≈ P × threads should roughly match physical cores; over-subscribing makes every process slower. A common sweet spot is a handful of processes each with a few threads (e.g. an 8-core box: -P 4 --threads 2, or -P 2 --threads 4). On a GPU box, run enough processes to keep the card busy but watch VRAM — each process loads its own copy of the model. Benchmark two or three (P, threads) points on a representative subset before committing to the full corpus; the benchmarking guide has the proof-of-work rules (a crash or wrong-model load can fake a “fast” run).

This scales linearly with cores/machines, needs no server, and each file’s success/failure is independent (easy to retry just the failures).

If you’d rather push files at a persistent server (model stays warm, no per-file process spawn), run it with --server-workers N and fan out client requests — but for a one-shot corpus the CLI fan-out above is usually simpler and at least as fast.


Horizontal scaling (replicas + load balancer)

For a serving deployment that must handle sustained concurrent traffic — or to scale past one box — run N independent server processes/containers behind a load balancer. This is the standard, robust pattern, and it is exactly the “8 replicas behind a proxy → several-× throughput” setup people arrive at by hand. It composes with (or replaces) the in-process --server-workers pool: each replica is a full server, so it also sidesteps the pure-ASR routing caveat and the shared-model /load restriction.

Docker Compose

The repo ships a docker-compose.yml with a service named crispasr. Scale it and put a proxy in front:

# N independent server containers, each with the model loaded once.
docker compose up --build --scale crispasr=4

Port gotcha. The shipped compose publishes one host port for the whole service ("${CRISPASR_HOST_PORT:-${CRISPASR_PORT:-8080}}:${CRISPASR_PORT:-8080}"8080:8080 unless you override it, and the same value for every replica). Scaling as-is makes the replicas collide on that host port (“port is already allocated”). The fix is to not publish the backend port on the host and instead let the load balancer reach the replicas over the compose network (as below) — remove the ports: mapping from the crispasr service and publish only the proxy. (Docker’s DNS round-robins the service name to all replica IPs; a real LB with health checks is better.)

Then front them with any load balancer. A minimal Caddy config (round-robin / least-conn over the replicas):

:8080 {
    reverse_proxy crispasr:8080 {
        lb_policy least_conn
        health_uri /health
    }
}

GET /health is public (no API key) precisely so a load balancer / orchestrator can health-check each replica. nginx upstream { least_conn; server … } works the same way.

In-process pool vs. replicas — which to use

  --server-workers N (in-process pool) N replicas + load balancer
Setup one flag compose --scale + a proxy
Memory N× weights in one process N× weights across N processes
Concurrency scope pure-ASR requests only (see caveats) every request, no routing caveat
Model hot-swap (/load) disabled while pooled per-replica (roll one at a time)
Fault isolation one process — a crash takes all workers a crashed replica is drained by the LB
Cross-machine no (single process) yes
Best for one GPU under-utilised by a single stream production serving, multi-box, mixed requests

Rule of thumb: one GPU, small model, want concurrency cheaply → --server-workers. Production serving / multiple machines / mixed request types → replicas behind a load balancer. They stack: e.g. 2 replicas × 2 workers.


What is not supported, and why

No batched multi-stream inference

There is no API that takes N different audio clips and runs them through one model instance as a single batched forward pass. “Batch” inside the codebase means splitting one long stream into sequential chunks with context carry — not cross-clip batching. Concurrency is achieved by running independent instances (workers / processes / replicas), each handling one stream, never by batching several streams on one instance.

Why: the backend zoo is heterogeneous (encoder-heavy CTC/transducer models, small AR decoders, flow-matching TTS, …) with per-model preprocessing, chunking, and decoding. A generic cross-stream batch path would have to be built and verified per backend, and for the dominant use case (offline bulk) file-level process parallelism already saturates the hardware with none of that complexity.

A single loaded context is not thread-safe

One crispasr_context / session must not be driven from multiple threads at once (documented in include/crispasr.h). Concurrency comes from separate instances, not shared-context threading. If you use the C-ABI / a binding directly, serialize calls per context or create one context per thread.

PagedAttention is not the right fit

PagedAttention (vLLM) solves KV-cache fragmentation and memory pressure when serving many concurrent, long, autoregressive LLM sequences that share a giant paged KV cache — it lets one model instance pack many sequences into VRAM efficiently.

That is not where CrispASR’s ASR throughput goes:

So the honest recommendation is the replica/worker model documented above, not a PagedAttention port. If a specific large-AR-decoder backend ever becomes the throughput bottleneck under concurrent serving, KV-cache paging could be revisited for that backend — but it would not be a general CrispASR feature.


See also