At 09:58, an API fleet starts missing its latency target. CPU is available, the database is healthy, and the load generator is already sending more requests than the service handled yesterday. The first instinct is often to search for the fastest web servers and replace the front door. That may help. It does not prove that the webserver is the bottleneck. Profile the request path first, then test whether the front door is limiting you. For the application-layer distinction, see what a web service actually is.
A deployment review can produce the opposite trap: the benchmark looks excellent, but the proposed stack cannot load one required component dynamically, cannot preserve a buffer across an asynchronous boundary, or spends its gains waiting on a downstream service. Use a benchmark to narrow the shortlist. Do not use it to make the workload decision for you.
- These are three high-throughput application/platform entries worth testing, not a universal top-three ranking: uWebSockets.js, ntex [tokio,platform], and ASP.NET Core [Platform, Pg, AOT].
- TechEmpower Round 23 reports different maxima for HTTP Plaintext and JSON. The values use different pipeline or concurrency settings, so they are not one composite score.
- Recommendation: choose by workload mechanics and operating constraints first, then reproduce the result with TLS, dependencies, observability, and realistic traffic.
How to Evaluate the Fastest Web Servers
The benchmark source is TechEmpower Round 23’s physical Citrine run. The machine was a ProLiant DL360 Gen10 Plus with an Intel Xeon Gold 6330 at 2.00 GHz, 56 cores, 64 GB of memory, and Mellanox ConnectX-6 40Gbps Ethernet. TechEmpower’s test overview describes Plaintext as a routing-focused test using HTTP/1.1 pipelining. JSON adds request and header parsing, object creation, JSON serialization, response headers, and request throughput.
The table reports each selected entry’s highest observed value in each named dimension. Plaintext maxima come from pipeline levels; JSON maxima come from ordinary concurrency levels. uWebSockets.js and ASP.NET Core are close in Plaintext throughput, while ntex trails them there and leads the three entries in JSON throughput. The ordering changes because request shape changes which work dominates.
These are workload-specific measurements, not a universal production ranking. The cited figures do not compare TLS, HTTP/2, HTTP/3, database calls, business logic, or tail latency. They are not WebSocket measurements either. Read the raw records and configuration alongside the table; the table is a summary, not a substitute.
| Selected stack | Round 23 Plaintext max (HTTP pipeline) | Round 23 JSON max (ordinary concurrency) |
|---|---|---|
| uWebSockets.js | 27,991,649 requests/second | 2,731,253 requests/second |
| ntex [tokio,platform] | 24,925,516 requests/second | 2,963,471 requests/second |
| ASP.NET Core [Platform, Pg, AOT] | 27,770,995 requests/second | 2,556,896 requests/second |
- uWebSockets.js: Plaintext maximum at HTTP/1.1 pipeline level 4096; JSON maximum at ordinary concurrency 512. Raw totals were 419874736 and 40968793 over 15 seconds.
- ntex [tokio,platform]: Plaintext maximum at pipeline level 1024; JSON maximum at ordinary concurrency 512. Raw totals were 373882736 and 44452071 over 15 seconds.
- ASP.NET Core [Platform, Pg, AOT]: Plaintext maximum at pipeline level 4096; JSON maximum at ordinary concurrency 512. Raw totals were 416564928 and 38353445 over 15 seconds.
For the raw data, see TechEmpower’s Round 23 physical-hardware JSON. The Round 23 hardware article explains the environment, and TechEmpower’s test overview defines the workload categories.
Three Stacks, Three Different Trade-Offs
µWebSockets.js (uWebSockets.js): a Node.js service with a native boundary
µWebSockets.js, written uWebSockets.js in the benchmark identity, is a Node.js binding around a C++ implementation. You keep Node’s application model. But the native networking that puts it on a high-throughput shortlist also means your team must own native-addon builds and deployment, along with debugging across the JavaScript/C++ boundary, rather than treating it like a pure-JavaScript server.
Hypothetical use case: imagine a realtime collaboration service where browsers maintain many long-lived connections, edits are fanned out through pub/sub, and a small HTTP API handles session setup and presence. Market-data fanout is another plausible case: a Node.js gateway receives binary-oriented updates and pushes them to subscribed clients. These are workload descriptions, not claims that the Round 23 HTTP results prove WebSocket performance.
The official uWebSockets.js declarations document send status, buffered bytes, configurable maximum backpressure, drop behavior, and a drain callback. Those details determine what happens when a subscriber falls behind. The service must also respect the documented lifetime of callback-owned ArrayBuffer data: retaining it across the first await or return is unsafe. The documentation recommends corking writes after an asynchronous boundary where appropriate.
Choose uWebSockets.js when the team already works comfortably in Node.js and is willing to own those constraints. A large Plaintext row is not enough reason by itself. If your service is dominated by database waits, complex authorization, or downstream calls, native HTTP throughput may not change the result. Without reproducible native builds and an explicit backpressure policy, the operational cost can outweigh the benefit.
ntex [tokio,platform]: Rust async composition
ntex offers a Rust model for HTTP services that need typed composition around asynchronous work. Its documentation covers HTTP/1.x and HTTP/2, middleware, streaming and pipelining, TLS, and WebSockets. The benchmark identity matters: the cited row is ntex [tokio,platform], not a generic result for every ntex runtime or feature set.
Hypothetical use case: consider a Rust telemetry ingestion service receiving batches from thousands of agents. It validates structured payloads, streams accepted records toward a queue, and exposes health and administrative endpoints. A service that relays chunks between authenticated clients is another reasonable fit. Typed request handling and explicit async composition can make either design easier to reason about; the benchmark gives you a reason to measure the routing layer.
Selection requires discipline. ntex has multiple runtime and backend choices, and the Round 23 results contain separate ntex keys. Do not transfer the Tokio/platform values to compio, database, or micro entries. Keep blocking work off the async execution path, because a CPU-heavy parser or synchronous client hidden inside a handler can erase the benefit of a fast network loop.
I would pick ntex when Rust’s ownership model and service-level control are part of the design, not as a language switch made for a table row. I would avoid a benchmark-led choice if your team lacks a credible Rust operating model or if the service spends most of its time in external systems. The benchmark does not measure staffing, ecosystem breadth, or production support, so it cannot settle those questions.
ASP.NET Core [Platform, Pg, AOT]: deployment efficiency with constraints
The ASP.NET Core entry uses Kestrel as its web server and is specifically labelled ASP.NET Core [Platform, Pg, AOT]. That qualifier matters. It is a result for one Platform/Pg/AOT implementation, not a throughput claim about every ASP.NET Core application or every Kestrel deployment.
Hypothetical use case: suppose a platform team operates a large fleet of small Minimal APIs generated from an internal service definition. Each service has a narrow contract, limited reflection, and a predictable set of dependencies. A self-contained native executable, startup behavior, and instance density may matter more than keeping every dynamic framework feature. A source-generator-friendly fleet of Minimal APIs is a more credible AOT candidate than a general MVC application with runtime discovery everywhere.
Microsoft’s ASP.NET Core Native AOT guidance says this is not a drop-in switch for all applications. A related native-versus-JIT comparison illustrates why startup and runtime trade-offs still need workload testing. In the cited compatibility view, Minimal APIs are partial, MVC is unsupported, SignalR is partial, and WebSockets are supported. Native AOT also disallows dynamic loading and runtime code generation, requires trimming and single-file compatibility, and surfaces publish-time warnings that need review. Test the published artifact, not just a development build.
Choose this path when the application fits those compatibility boundaries and fleet economics, startup, or self-contained deployment are real constraints. Do not generalize the AOT caveats to all Kestrel. Do not adopt AOT because a benchmark table contains “AOT” as a label. First prove that the framework features your service uses survive publishing, trimming, observability setup, and upgrades.
Where Nginx, Caddy, Lithium, and WebSockets Fit
Nginx and Caddy belong in a different layer. Their official documentation describes reverse-proxy behavior: they accept traffic at the edge and forward requests to upstream applications. An Nginx reverse proxy or Caddy’s reverse_proxy directive can handle TLS termination, routing, certificates, and operational policy. Neither replaces the application/platform rows above, and no Nginx or Caddy deployment is measured in the six values.
Lithium is a relevant comparator in the raw results. Its official repository describes a C++17 HTTP server and publishes its own project information. Individual Lithium tests can exceed some listed rows. A strict numeric ranking still requires a declared metric, test dimension, level, and inclusion set, so this article does not create one.
The same separation applies to WebSockets. uWebSockets.js and ntex document WebSocket capabilities, but the cited values are HTTP Plaintext and JSON results, not WebSocket measurements. Use the documentation to form a use-case hypothesis, then test connection churn, message size, fanout, backpressure, and tail latency with your own traffic.
A Practical Selection Guide
Start with compatibility and the operating model. For the fastest web servers in a real service, fit usually matters more than a headline maximum. Pick uWebSockets.js when Node.js integration or realtime connections matter and the team accepts native-addon ownership. Binary-oriented messaging is another reason to consider it. Pick ntex [tokio,platform] when Rust’s type and async model fit the service. Its middleware composition, streaming, and protocol control still need to match the design. Pick ASP.NET Core [Platform, Pg, AOT] when a source-generator-friendly application fits Native AOT and fleet efficiency matters.
Add Nginx or Caddy when you need an edge component. That proxy choice does not replace selection of the application stack. Platform teams making that boundary decision may also find this platform-engineering comparison useful context. Reject benchmark-led selection when database latency, downstream calls, serialization complexity, authorization, or business logic dominate the endpoint.
My preferred process has two stages: eliminate incompatible stacks first, then benchmark the survivors with production-shaped traffic. This keeps a large routing number from hiding a poor fit.
Production Testing Checklist
- Record the exact stack, version, compiler or runtime, entry and feature set, operating system, CPU topology, memory, NIC, kernel, and container image.
- Reproduce the request shape, payload size, connection behavior, concurrency, pipeline setting, and duration. Keep the raw output and configuration.
- Add TLS, the HTTP versions, and the proxy path used in production. Do not infer encrypted or HTTP/2/HTTP/3 behavior from these plaintext figures.
- Test realistic JSON schemas, headers, authentication, authorization, compression, logging, tracing, metrics, and error responses.
- Include databases, caches, downstream APIs, queues, and representative business logic. Measure the complete service path.
- Measure throughput, median latency, p95, p99, p99.9, errors, CPU, memory, allocator or garbage-collection behavior, connection counts, and saturation points.
- Run cold-start and warm-up tests, especially for Native AOT and autoscaling. Inspect the published artifact.
- For uWebSockets.js, test backpressure, buffered bytes, drain behavior, and buffer lifetime across async boundaries.
- For ntex, test the selected runtime/backend and isolate blocking work from async execution.
- For ASP.NET Core AOT, inspect publish warnings, trimming behavior, compatibility gaps, and every framework feature the service uses.
- Repeat the runs and report variance. Close maxima are close results, not proof of a winner.
Key Takeaways
- uWebSockets.js is a candidate for Node.js services where realtime behavior and native networking justify explicit buffer and backpressure management.
- ntex [tokio,platform] is a candidate for Rust services where typed async composition and streaming control fit the team and workload.
- ASP.NET Core [Platform, Pg, AOT] is a candidate for compatible Minimal API fleets where deployment efficiency matters; do not generalize that result to all Kestrel.
- The Round 23 figures are useful starting points, not a universal fastest web server verdict. Your endpoint, dependencies, security controls, and latency objectives decide the outcome.
Conclusion & Next Steps
These three stacks are worth testing under the right conditions, but none is among the fastest web servers for every workload. The benchmark compares specific entries and workload dimensions: uWebSockets.js, ntex [tokio,platform], and ASP.NET Core [Platform, Pg, AOT]. It does not measure your database, TLS path, business logic, WebSocket traffic, or tail latency.
Choose one representative endpoint. Reproduce the baseline. Add production dependencies and security controls. Run the checklist, then record what changed your decision. If you test one of these stacks, share the workload, conditions, and failure mode, not just the headline number.
Frequently Asked Questions
What are the fastest web servers?
There is no universal answer. These Round 23 entries are high-throughput candidates under specific workloads, not a production ranking for every service.
Do these figures measure WebSocket performance?
No. The cited values are HTTP Plaintext and JSON results. WebSocket support must be tested with your connection, message, fanout, backpressure, and latency profile.
How should I choose among these stacks?
Filter by runtime, compatibility, operating model, and deployment constraints first. Then benchmark a representative endpoint with production dependencies, security controls, and realistic traffic.
Sources: TechEmpower environment; pinned pipeline runner; ntex documentation; .NET Native AOT deployment guidance.
