write_request_line checked the request target for CR/LF but concatenated
the method verbatim. A method carrying CR/LF could smuggle a whole
request ahead of the real one, and the client would take the smuggled
request's response as its own. A method with a space or an empty method
put a malformed request line on the wire.
Require the method to be a token (RFC 9110 Section 9.1) before anything
is written. All three callers (the buffered client path, open_stream and
the WebSocket handshake) go through this function and fail with
Error::Write, as they already do for a rejected target.
Claude-Session: https://claude.ai/code/session_01NTDesJQTQPEuu4o4XCu69g
* reject control characters in chunk extensions in read_payload
* Bound every chunk-size line scan by the line terminator
read_payload() ended its scans of one line buffer two different ways: the
hex-size parse and the space skip that follows stopped on the NUL that
stream_line_reader::append() writes, while the new chunk-ext check walked
to an explicit end pointer. Compute that end pointer first and bound all
of them by it, so no scan depends on the buffer's NUL and the terminator
can never be read as line content.
The bare-LF branch is reachable only under
CPPHTTPLIB_ALLOW_LF_AS_LINE_TERMINATOR, where getline() ends the line on
an LF that is the terminator rather than extension text. Say so: the
comment below it explains why a bare LF inside the line is rejected, and
without that note the two read as contradictory. Its guard no longer
depends on the scan cursor either, since all it ever needed was a check
that there is a byte to look at.
* Reuse the chunked-body helper in the chunk-ext acceptance test
AcceptsChunkExtension repeated expect_chunked_body_rejected()'s body
verbatim apart from the expected status, so parameterise the helper on
the status and keep the rejection wrapper for the existing callers. The
decoded body is already checked by the /chunked handler, so asserting
the status is all the new test needs.
Also record why the control-character literal stays split: a hex escape
consumes every hex digit that follows it, so "\x01b" would be the single
byte \x1b rather than \x01 followed by 'b', and joining the halves would
quietly change what the test sends.
---------
Co-authored-by: yhirose <yuji.hirose.bug@gmail.com>
tls::shutdown() on OpenSSL called SSL_shutdown() a second time to wait
for the peer's close_notify. An idle keep-alive client never sends one,
so closing its connection held the worker thread until the read timeout,
and Server::stop() waited for it. Send close_notify and return, as the
Mbed TLS and wolfSSL backends already do.
The server wrote 100 Continue as soon as it saw the expectation, before
pre_routing_handler, pre_request_handler, or routing ran. A request
those handlers rejected, or one that matched no route, still invited
the client to send a body the server would never read.
Defer the interim response until the body is about to be read. If the
request is answered without reading the body, 100 Continue is never
sent and the connection is closed, since whether and when the client
sends the body is unknown.
Also treat a 417 returned by expect_100_continue_handler as the final
response. It used to be written as a bare status line, after which the
request was processed and a second response was written.
The WebSocket upgrade path matched the route and switched protocols
without setting req.matched_route or calling pre_request_handler, so a
check placed there (e.g. authentication) never ran for WebSocket
routes. Set matched_route and run the handler before the upgrade; if it
handles the request, reply with a regular HTTP response instead of 101.
Also write rejected upgrade responses (from pre_routing_handler too)
with write_response_with_content, so they carry Content-Length. Without
it, a client reading the body waited until the keep-alive timeout.
`sed -i ''` is BSD-only: GNU sed takes the '' as the script and the expression as a file name, so `just release --run` failed on Linux before touching anything.
199d7ee made read() return the new ReadResult::Timeout for every read
timeout and leave the connection open. The compile-time server default
(CPPHTTPLIB_WEBSOCKET_SERVER_READ_TIMEOUT_SECOND, 300s) is always in
effect, so a handler written as `while (ws.read(msg))`, the form the
README's Quick Start uses, no longer ended when a peer went quiet:
Timeout is non-zero, so the loop ran its body again with the previous
message still in `msg`, and the worker the backstop is meant to reclaim
was never released. Nothing caught it because every test of the new
result set a timeout explicitly and checked the result by value, and the
heartbeat tests keep the connection alive with pings.
The two timeouts mean different things. One the caller sets through
set_read_timeout() is a request for control back, and is reported as
Timeout on a still-open connection. The compile-time default is a
backstop against a peer that has gone quiet, and elapsing it is now a
failure again: read() returns Fail and closes the connection, as it did
before 199d7ee. WebSocket tracks whether set_read_timeout() was called,
and WebSocketClient carries the same flag over to the WebSocket it
creates on connect().
Tests use the heartbeat binary, which compiles both defaults down to 3s:
a `while (ws.read(msg))` server handler runs its body once and exits
when the client falls silent, and a client that never set a timeout gets
Fail with the connection closed. The README and cookbook now say which
timeout produces Timeout.
Claude-Session: https://claude.ai/code/session_01EF5uZ1X2kaHhqJ8VgfjVaQ
ProxyTest, RedirectTest.HTTPBin*, KeepAliveTest and ProxyTest.SSLOpenStream
still sent their requests through the squid proxies to the external
httpbingo.org, so an upstream hiccup there failed CI with no code change
involved (KeepAliveTest.SSLWithDigest got a 502 on its first /get).
Switch them to the "httpbin" container (nginx + go-httpbin) that
BaseAuthTest/DigestAuthTest already use. go-httpbin serves /get,
/redirect/n and /digest-auth the same way, so the test logic is
unchanged; the SSL variants disable certificate verification for the
self-signed test cert, as BaseAuthTest.SSL does.
RedirectTest.YouTube* is left pointing at youtube.com since it exercises
a real cross-host, http -> https redirect chain.
Claude-Session: https://claude.ai/code/session_0148ZAzsuYRYXkwcA7UFeh95
The intermittent failures it was tracking stopped after the graceful
drain before close in 8e702d3: no windows-without-SSL test failure on
master since 2026-08-09. Drop the reporting step, the issues: write
permission it needed, and the run_tests step id that only it used.
Claude-Session: https://claude.ai/code/session_0148ZAzsuYRYXkwcA7UFeh95
* reject ambiguously framed responses in client read paths
* Accept non-chunked Transfer-Encoding responses in the client framing guard
RFC 9112 §6.3 treats requests and responses differently when the final
transfer coding is not chunked: a request's body length cannot be
determined and the server must answer 400, but a response's body simply
runs until the server closes the connection. read_content() and the
open_stream() body reader already do that, so such a response is not
ambiguous and rejecting it broke valid responses such as
"Transfer-Encoding: gzip" followed by a close.
Keep rejecting a Transfer-Encoding paired with a non-zero Content-Length,
which is the actual ambiguity, and drop the non-chunked clause from both
client read paths.
Tests: check that rejection surfaces as Error::Read, that a non-chunked
Transfer-Encoding response is read until close on both paths, and that
HEAD, 204 and 304 responses with both framing headers are not rejected.
Claude-Session: https://claude.ai/code/session_01JYPWKpbp4a881EdpEf2xSi
* Share the framing check and reuse existing test helpers
Factor "Transfer-Encoding with a non-zero Content-Length" into
detail::has_conflicting_content_length() next to
is_chunked_transfer_encoding(), and call it from the server request
guard and both client read paths so the rule and its RFC 9112 §6.3
rationale live in one place.
In the tests, drop the POSIX-only raw socket helper in favour of the
existing serve_single_response() and read_all(), which also lets the
tests run on Windows. Fold the stream-only test into the buffered one so
each case checks both Get() and open_stream(), and cover the HEAD/204/304
exclusion on the open_stream() path too.
Claude-Session: https://claude.ai/code/session_01JYPWKpbp4a881EdpEf2xSi
---------
Co-authored-by: yhirose <yuji.hirose.bug@gmail.com>
* Reject a Range first-byte-pos that overflows ssize_t
parse_range_header initializes first to the -1 sentinel that means "no
first-byte-pos" and only overwrites it when detail::from_chars succeeds.
On std::errc::result_out_of_range the assignment is skipped and -1
survives, so "bytes=9223372036854775808-100" is parsed as the suffix
range "bytes=-100" and range_error serves the last 100 bytes instead of
returning 416.
Before the parser was rewritten onto detail::from_chars, std::stoll threw
std::out_of_range on the same input, the catch arm added in 8f8761e for
issue #705 returned false, and the request was answered with 416. The
catch arm is still there but from_chars reports through an error code, so
nothing reaches it any more.
get_header_value_u64 and parse_port already reject an out-of-range value
at their from_chars call sites; this was the remaining one that dropped
the error.
The last-byte-pos side is deliberately unchanged: -1 there is the
documented RFC 9110 14.1.2 "remainder of the representation" value, so an
oversized last-byte-pos stays accepted.
* Simplify the Range first-byte-pos overflow check
Parse the first-byte-pos straight into first, since a failed parse now
returns before first is read, and fold the overflow test into the
existing batch of rejected ranges. Also note on the last-byte-pos side
why an overflow there deliberately keeps -1.
Claude-Session: https://claude.ai/code/session_01JYPWKpbp4a881EdpEf2xSi
---------
Co-authored-by: yhirose <yuji.hirose.bug@gmail.com>
* send each credential only to its own hop in write_request
An SSLClient behind a proxy sent Proxy-Authorization inside the TLS tunnel, where the origin reads it, and sent the origin's Authorization on the CONNECT request the proxy reads. Attach each only on the message its hop actually reads.
* Keep default headers off the CONNECT request
set_default_headers() is typically used for origin credentials such as
Authorization, Cookie or API keys, but they were also attached to the
CONNECT request an SSLClient sends to its proxy, in plaintext before the
TLS tunnel exists. Default headers now go only on requests the origin
reads, the same split the previous commit makes for set_basic_auth and
set_bearer_token_auth.
Claude-Session: https://claude.ai/code/session_01JYPWKpbp4a881EdpEf2xSi
* Simplify per-hop credential handling and its tests
Flatten the Authorization insertion in write_request into one guard with
an else-if (Basic already took precedence over Bearer), and shorten the
comments around it. Fold DefaultHeadersStayOffConnect into the
CredentialsStayWithTheirHop helper, which now takes the list of headers
that must reach only the origin.
Claude-Session: https://claude.ai/code/session_01JYPWKpbp4a881EdpEf2xSi
---------
Co-authored-by: yhirose <yuji.hirose.bug@gmail.com>
These tests exercise the squid proxies by hitting /basic-auth and
/digest-auth on an external httpbin-style site. That site's identity has
already moved twice (httpbin.org -> httpcan.org, per #2300) chasing
uptime, and httpcan.org itself is now down (Cloudflare 502 from its
origin), failing CI with no code change involved.
Adds two containers to the existing squid docker-compose stack instead:
go-httpbin (mccutchen/go-httpbin) as the backend, and an nginx sidecar in
front of it under the single "httpbin" hostname so both the NoSSL tests
(port 80) and the SSL tests, which CONNECT-tunnel through the proxy to
port 443, resolve the same name -- go-httpbin only listens on one port at
a time, so it can't serve both protocols itself. nginx uses the repo's
existing self-signed test cert; the SSL client tests already disable
verification for it like other self-signed-cert tests in this suite.
go-httpbin was picked over the more feature-complete kennethreitz/httpbin
after finding the latter accepts a wrong digest-auth username as long as
the password matches -- confirmed with a direct curl against the
container, unrelated to anything in this repo. go-httpbin correctly
rejects both. The trade-off is losing SHA-512 digest-auth coverage here,
since go-httpbin only implements MD5 and SHA-256; nothing else in the
suite exercises SHA-512 digest auth against a live server. Response body
assertions are adjusted to go-httpbin's actual JSON shape (an added
"authorized" field, no "algorithm" field), and the domain changes from
httpcan.org to the self-hosted "httpbin".
This only affects 'make proxy'/'make proxy_mbedtls'/'make proxy_wolfssl'
and the Proxy Test CI workflow -- the default 'make' target is untouched.
The previous commit pinned CI and the pre-commit hook to a fixed
clang-format 23.1.0, but the maintainer develops on macOS against
whatever version `brew` currently installs, which changes over time
as Homebrew updates the formula.
style-check now runs on macos-latest and installs clang-format via
`brew install`, so it tracks the same moving target the maintainer's
Mac does. The pre-commit hook switches from pre-commit's own pinned
mirror to a local hook that shells out to the system clang-format,
so a local commit and CI both go through the same Homebrew-installed
binary rather than two independently versioned copies.
Trade-off: this reintroduces the non-determinism a fixed pin avoids
-- a commit's style-check result can now change over time as Homebrew
updates the formula -- but that mirrors how the maintainer already
develops, which is the point.
Also install coreutils in CI: the style_check Makefile target needs
grealpath's --relative-to, which the macOS-native realpath lacks.
CI relied on ubuntu-latest's default apt clang-format (18.1.3), while the
pre-commit hook was pinned to a different 18.x build. Neither tracked a
specific, deliberately-chosen version, and the two could drift from each
other and from whatever a contributor has installed locally.
Both now pin the same clang-format 23.1.0 (latest stable), installed via
pipx in CI. Also scope the pre-commit hook's file matcher to the same set
of files test/Makefile's style_check target checks, since the previous
\.(cpp|cc|h)$ pattern reached into vendored code (test/gtest,
benchmark/crow) that must stay untouched.
Reformat httplib.h's brace-init spacing to match 23.1.0's output.
Defined next to its non-template overload, it landed in the part of the header
that test_split compiles into httplib.cc, leaving a user TU's instantiation
with nothing to link against. The WebSocketClient overload of the same name has
always sat up with the class definitions; this one belongs there too.
Only test_split sees it -- a header-only build instantiates the template
wherever it is written, which is why the regular test target stayed green.
read() collapsed every failure into Fail and marked the connection closed with
it, so a read timeout could not be used to get control back and send on the
same connection -- it killed the connection instead. The information was
already there and thrown away: SocketStream::read records Error::Timeout, and
read_websocket_frame flattened it into a bool.
ReadResult gains Timeout, reported only when the timeout elapsed on a frame
boundary with nothing consumed, which is the only case where the stream can be
read again. Every multi-byte field now loops until it has its bytes, which also
fixes a frame header straddling the read buffer's boundary failing the frame:
Stream::read is allowed to return less than asked for, and only the payload
was reading in a loop.
ws::WebSocket::set_read_timeout() lets a server handler bound its own reads,
and WebSocketClient::set_read_timeout() now reaches an already-open connection
instead of only seeding the next connect().
The read timeout macro splits in two. A client waits forever by default -- a
read timeout is the caller's tool for taking back control, not a liveness
check, which is ping/pong's job -- while a server keeps the 300s that reclaims
a worker from a peer gone quiet. Defining the old name still sets both.
Also record a reason on the two WebSocketSSLStream::read failure paths that
returned -1 without one, so get_error() cannot report a previous call's
timeout, and make SocketStream's read timeout atomic now that it can be
changed while a read is in flight.
* Don't compress a response whose handler already set Content-Encoding
* Don't compress a pre-encoded response served from a file
The guard that stands down when a response already names a content coding
covered the responses that settle their coding in `apply_ranges()`, but a
file-backed one settles it in `static_file_encoding()`, which asked the
content-type overload and so never saw the field. With static file
compression enabled, a mount point naming the coding for a tree of
build-time compressed assets, and a handler setting the field on a
`set_file_content()` response, both had their stored bytes compressed a
second time and a second `Content-Encoding` field line appended.
A file-backed response has not been given a content type by the time its
coding is decided, which is the only reason it could not go through
`encoding_type()`. It takes the type as an argument now, so both paths share
the one guard instead of carrying a copy each.
`Response::content_encoding_` becomes `content_coding_`, after what it
holds. It names the coding chosen for the body, which is what its own
comment already called it, while the old name read as the value of the
`Content-Encoding` field whose presence is exactly what forces the coding to
`None`.
README gains the behaviour, including the part that stays with the handler:
`Vary` is added only to a coding the server chose, so a handler that picks a
representation from `Accept-Encoding` has to add the field itself.
---------
Co-authored-by: yhirose <yuji.hirose.bug@gmail.com>
parse_www_authenticate() accepted any WWW-Authenticate: Digest
challenge that carried at least one auth-param, so a server sending
e.g. Digest qop="auth" with no realm/nonce would make it through.
make_digest_authentication_header() then dereferences auth.at("realm")
and auth.at("nonce") unconditionally, throwing std::out_of_range with
no try/catch on the retry path, which terminates the client process.
Now require both realm and nonce before treating a Digest challenge as
usable, same as if no Digest challenge were present at all.
The reporting step for issue #2533 was gated on failure(), which is true
when any step in the job failed. Two build failures on feature branches
were posted to the issue as flaky test recurrences, with the body falling
back to "Could not extract failed test name" because no test had run.
Gate it on the test step's own conclusion instead, and skip posting when
no [ FAILED ] line turns up in the shard logs. The explicit failure()
stays because an if expression with no status check function gets an
implicit success().
CPPHTTPLIB_MULTIPART_FORM_DATA_FILE_MAX_COUNT is enforced only in
Server::read_content(), where parts are accumulated into req.form. The
streaming ContentReader path keeps nothing and was never in scope, but
this was undocumented (GHSA-923p-8q8g-xcqj). Note the split in the README
and show how to bound the part count from inside a ContentReader handler.
* Drop the claim that small bodies skip compression
There is no size threshold anywhere in the compression path.
encoding_type() gates on the content type and Accept-Encoding only, and
apply_ranges() compresses whatever body it is given, so a two-byte
text/plain response comes back gzipped at 22 bytes.
Say what actually happens and leave the decision to the handler.
* Compress static file responses behind an opt-in (Fix#2545)
apply_ranges() runs the compressor inside the branch it takes when
res.body is non-empty. A response served from a file leaves res.body
empty and sets content_length_, so it took the other branch, which
writes Content-Length and returns; encoding_type() was computed before
the split and never consulted on that side. The same bytes handed to
set_content() came back gzipped, which left set_mount_point() and
Response::set_file_content() as the one path that missed out.
Add Server::set_static_file_compression(), off by default so nothing
about an existing server changes. When it is on, the file-backed
provider is run through the compressor into res.body ahead of the rest
of apply_ranges(), so the response is framed the way set_content()
already frames one: it keeps its Content-Length, and HEAD still reports
the size a GET would return.
Ranges are answered from the identity representation, since RFC 9110
applies Range after content coding and slicing a compressed body would
mean compressing the whole file first. The ETag carries the coding it
belongs to, so a client that cached the compressed form revalidates
against its own validator rather than the identity one. Both the ETag
and the body take their coding from static_file_encoding(), so the two
cannot disagree.
Providers registered with set_content_provider() are left alone. zlib
buffers until its window fills, so running one through a compressor
would hold back writes that a caller expects to reach the peer as they
are produced.
The compressed bytes stay in memory until the response has been
written, so the peak cost scales with requests in flight.
set_static_file_compression_max_length() bounds it, defaulting to 4MB.
* Add a minimum size for static file compression
Compressing a file that already fits in a single 1500-byte MTU does not
get it to the client any sooner, and a file of a few bytes comes back
larger than it went in once gzip's header and trailer are added. Every
other server draws this line: nginx's gzip_min_length, Caddy's
minimum_length, IIS's minFileSizeForComp, CloudFront's 1000-byte floor.
The note this replaces told callers to decide in the handler. A response
served through set_mount_point() has no handler to decide in, so the
floor has to live in the server. It defaults to 1400 bytes, the size
that fits inside one MTU with room for headers.
set_static_file_compression_min_length() moves it, and
CPPHTTPLIB_STATIC_FILE_COMPRESSION_MIN_LENGTH sets the default at
compile time. The empty-file case keeps its own early-out so that a zero
floor still cannot turn an empty body into a 20-byte gzip stream.
The two bounds now read as a pair, so the documentation says what each
one is for: the lower bound is about what is worth compressing, the
upper bound about what one request is allowed to cost.
Every file under test/www except 1MB.txt is below the default floor, so
the tests that need a small file compressed lower it explicitly.
parse_disposition_params() and extract_media_type() both split on every
';' and then on every '=', with no idea that a parameter value can be a
quoted-string. RFC 9110 5.6.6 allows ';' and '=' inside one, so
filename="report=v2.pdf" came out as v2.pdf", and filename="a;b.txt" was
truncated at the semicolon and left a bogus parameter behind.
The same defect reached the boundary. RFC 2046 5.1.1 allows '=' in a
boundary, which forces a sender to quote it, so the common MIME form
boundary="----=_NextPart_000_0000_01D9" parsed as
_NextPart_000_0000_01D9".
Add split_unquoted(), which is split() with the one extra rule that a
delimiter inside a quoted-string is not a delimiter, and route both
parameter parsers through it. The key/value split, duplicated verbatim
in the two of them, moves into divide_param_pair(). That one divides at
the first '=' without tracking quotes: 5.6.6 makes the key a token, so
no quote can precede the separator, and reusing divide() keeps this off
the per-byte scan.
A backslash stays an ordinary character here. Both browsers and
httplib's own sender percent-encode '"' rather than escaping it, and
recognizing a quoted-pair without also unescaping it would just trade
one wrong value for another.
write_content_with_progress() advances its offset only by what the provider
writes, so a provider that reported success without writing anything and
without calling done() was handed the same offset and length again on the next
pass. With the peer still connected it spun there, re-entering the provider as
fast as the loop could run.
make_file_body()'s provider was one way to reach this and was fixed in #2566,
but any user-supplied provider can do the same. Treat a pass that makes no
progress as a short body, which is how done() called early is already handled.
make_file_body() measures the file once and that length is already the
response's Content-Length. The provider re-opens the file by path on each
call, so if the file has been truncated since, the read comes up empty and
the provider returned true without writing. write_content_with_progress()
advances its offset only by what was written, so it called the provider
again, got nothing again, and kept spinning until the peer gave up.
Return false instead, as every other failure in this provider does.
parse_accept_header() rejected any Accept value with a leading, trailing
or doubled comma, and Server::process_request() validates Accept before
routing, so "Accept: text/html," was answered 400 Bad Request on every
route.
RFC 9110 Section 5.6.1.2 requires a recipient to parse and ignore empty
list elements in a #rule list, so those values are legal. split() already
trims each element and skips the empty ones, which made the guard inside
the callback unreachable as well; drop both and let the empty elements
fall away. The header length limit bounds how many a sender can send, so
ignoring all of them cannot be used as a denial-of-service vector.
get_combined_header_value() keeps skipping empty field lines, but that
skip is no longer observable through a request now that a stray comma
parses cleanly, so it gets its own test.
Server::process_request() wraps only routing() in a try/catch.
Everything else the user supplies runs outside it:
- the content provider, from write_response_core()
- post_routing_handler_, error_handler_, logger_
- expect_100_continue_handler_
- a WebSocket handler, and pre_routing_handler_ on the upgrade path
An exception from any of those unwinds out of process_and_close_socket()
into the task queue, which calls the job without a catch, so it reaches
the top of a pool thread and terminates the process. One handler that
throws takes down every other connection the server is holding.
Add Server::serve_guarded() and run the serving loop through it in both
process_and_close_socket() overloads. The exception is not turned into a
500: by the time a content provider runs, the status line and headers
are already on the wire, so there is nothing left to replace. Report it
through the error logger as Error::UserCallbackException and drop the
connection, which is what the peer observes regardless. Requests on
other connections are unaffected, and the socket is still drained and
closed - which unwinding used to skip on the non-SSL path, since
drain_and_close_socket() sits after the call rather than in a scope
guard.
The error logger is a user callback too, so the report inside the guard
is itself wrapped: a throwing logger must not be able to open the guard
back up.
Adds ServerExceptionTest: a throwing content provider, post-routing
handler, WebSocket handler and error logger, plus the content provider
case against SSLServer, each checking that a later request on a new
connection still succeeds. Every test runs the server on a single worker
thread, so a guard that catches the exception but still loses the thread
shows up as the follow-up request never being served. Note that all of
them abort the test binary without this change - which is the bug, but
it means a regression here fails the run rather than one test.
write_content_chunked()'s sink treated "the provider wrote nothing" as
"the provider has finished":
data_available = l > 0;
so sink.write(p, 0) ended the loop. Only done()/done_with_trailer()
emit the terminating zero-length chunk, so the body was left
unterminated - and the function still returned Success, because the
post-loop check only reports the is_shutting_down() case. The peer waits
for a last chunk that never arrives, and on a keep-alive connection
anything written next is parsed as a chunk-size line.
A provider reaching a pass with nothing to hand over is ordinary:
popping an empty buffer off a queue, or a compressor that has consumed
its input without producing output yet. It is not the end of the
message.
Ignore zero-length writes instead. A zero-length chunk is the terminator
in chunked coding, so it must never be emitted mid-body either way, and
data_available is now controlled only by done()/done_with_trailer().
This matches write_content_without_length(), where the sink's write
never ends the body.
The old behaviour cannot have been relied on: it produced an
unterminated response, so a provider using it never worked in the first
place.
DataSink has four callbacks, but only write is assigned by every writer
that hands a sink to a content provider:
write_content_with_progress() write, is_writable
write_content_without_length() write, is_writable, done
write_content_chunked() all four
send_with_content_provider...() write
get_multipart_content_provider() write, done (cur_sink)
A provider that calls one of the unassigned ones invokes an empty
std::function and throws std::bad_function_call. Nothing on that path
catches it, so it unwinds out of the thread running the provider and
terminates the process. The README's own idiom is enough to hit it:
sink.done() is documented for the without-length overload, but a
provider registered through set_content_provider() with a length gets a
sink where done is empty.
Default the three optional callbacks instead. A sink is writable unless
a writer says otherwise, and a sink that cannot carry trailers still has
to finish, so done_with_trailer() falls back to done(). Capturing this
for that is safe because DataSink is neither copyable nor movable.
A no-op done() alone would only trade the crash for a hang on the two
length-framed paths: both loop until offset reaches the promised length,
so a provider that reports itself done without writing would be called
again immediately, forever. Both now record that the provider finished
and stop, and the short body is reported as a write error. The client
path gains that check for the compressor-failure exit as well, which
used to send a truncated request body without reporting anything.
cur_sink in get_multipart_content_provider() now forwards is_writable
from the outer sink, so a provider item asking whether it may keep going
gets the stream's answer rather than the default.
The accept loop in Server::listen_internal() classified accept() failures
by reading errno, but Winsock reports them through WSAGetLastError() and
never touches the CRT errno. Both retry branches were therefore dead code
on Windows, and every accept() failure fell through to the fatal path,
which closes the listening socket and ends listen().
That is reachable in normal operation: a peer resetting a pending
connection before it is accepted is enough, and descriptor or buffer
exhaustion shows up under load. One such event stopped the server from
accepting anything again.
Add is_accept_resource_error() and is_accept_transient_error() next to
is_connection_error(), which already abstracts the same errno vs
WSAGetLastError() difference, and use them in the accept loop.
The POSIX sets are widened to match the Windows ones rather than being
left as they were: ECONNABORTED is the POSIX spelling of the aborted
pending connection that motivates this, and ENFILE, ENOBUFS and ENOMEM
are resource exhaustion in the same sense as EMFILE.
When accept() failed for a reason the retry branches do not cover, the
loop closed svr_sock_ but left the descriptor in the atomic. Two things
go wrong from there:
- A later stop() reads the stale value and calls shutdown()/close() on
it. By then the OS may have reused the descriptor for an unrelated
socket (a worker's keep-alive connection, or one the application
opened), and that connection is torn down instead.
- keep_alive() in the worker threads watches svr_sock_ to notice that
the server is going away, so the workers keep waiting on a listening
socket that no longer exists.
Take the descriptor with exchange(INVALID_SOCKET) before closing it,
which is what stop() already does. That also settles the race with a
concurrent stop(): whichever side takes the descriptor closes it exactly
once, and the other sees INVALID_SOCKET and does nothing.
parse_multipart_boundary only rejected an empty boundary, so a request could
declare one as long as a header line is allowed to be. A stock server accepts
up to 8146 bytes there, which is what CPPHTTPLIB_HEADER_MAX_LENGTH leaves after
"Content-Type: multipart/form-data; boundary=".
FormDataParser searches the body for "--" + boundary + CRLF with a plain
substring scan. buf_find scans for that delimiter's first byte, always '-', and
at every position that matches calls start_with, which compares until the first
mismatch. A body of '-' makes every position a candidate, and a boundary of '-'
makes each candidate compare the whole delimiter before failing at the CRLF. The
worst case is the product of the body length and the boundary length, and only
the first factor was bounded.
Measured by driving the parser directly in 16 KB reads, Apple clang 17 at
-O2 -DNDEBUG, best of three runs on an otherwise idle machine. 100 MB of '-',
the default payload limit, costs 2.59 s of CPU with a 70 byte boundary and
281.83 s with an 8147 byte one, a factor of 109. The same shape shows at 8 MB:
0.211 s, 3.081 s, 11.359 s and 22.091 s for boundaries of 70, 1024, 4096 and
8147 bytes.
RFC 2046 5.1.1 caps a boundary at 70 characters, so honoring that limit bounds
the multiplier too. The limit applies to the value after unquoting, so a quoted
70 character boundary stays valid. Only the server receive path parses a
boundary out of a Content-Type, so what clients may send is unaffected, and the
boundaries the library generates itself are 45 characters.
expect_split_multipart_ok() carries a comment saying the request sends
"Connection: close" so the response drain ends as soon as the server has
answered, but the header itself never made it into the request, so both
callers kept idling until the read timeout instead.
Add the header. EpilogueSplitAcrossReadsIsIgnored and
InitialBoundarySplitAfterLongPreamble each drop from about 3.1s to about
0.11s.
* Bound the multipart parser's buffer while it waits for a boundary
FormDataParser accumulated the entire request body whenever the declared
boundary never appeared in it. State 0 returned without erasing anything, so
the buffer grew to the full payload (100 MB by default) and buf_find rescanned
all of it on every 16 KB read. The cost grew with the square of the body size:
50 MB of '-' took 198 s of CPU on one core, and the buffer pinned the body in
memory for the whole request. One unauthenticated request was enough, and the
parser runs for any multipart request even when the handler never looks at the
parsed result.
State 0 now keeps only the last dash_boundary_crlf_.size() - 1 bytes while it
waits, which bounds both the memory and the rescan without capping how long a
preamble may be. The same 50 MB body now takes 0.14 s and the buffer stays at
one read plus the boundary. A boundary split across reads still parses, which
is what de5a255 (#2159) gave up this erase for.
State 4 buffered without bound in the same way when a boundary was followed by
neither CRLF nor "--". No further data can make such a body valid, so it now
fails right away. That is only safe because the close-delimiter branch moves to
a new state 5 that discards the epilogue: it used to stay in state 4, so an
epilogue arriving in a later read fell into this same branch. An epilogue
beginning with CRLF was then parsed as a new part and the request was rejected
with 400, which state 5 fixes as well.
Affected since v0.23.0, where de5a255 replaced the erase that had kept the
buffer in check.
* Skip buffering the multipart epilogue
Once the close delimiter has been parsed the parser is in state 5 and discards
whatever follows, but it still copied each epilogue read into the buffer before
erasing it. Return before buffering so a large epilogue spread across several
reads is dropped without being copied in at all.
* Clean up the multipart parser tests and the state 4 branch
Review follow-ups on top of the previous two commits, no behavior change.
- Move the four new tests next to the rest of MultipartFormDataTest. They
had landed in the middle of the RedirectTest block.
- Use bind_to_any_port instead of the fixed PORT, as AGENTS.md requires for
newly added servers. NoInitialBoundaryParsingIsNotQuadratic holds its port
for a couple of seconds, which matters when the suite is run sharded.
- Send "Connection: close" from expect_split_multipart_ok. The server kept
the connection alive after answering, so the response drain idled until the
client read timeout; both tests drop from about 3s to about 0.11s.
- Drop the dead `dash_.size() > buf_size()` guard in state 4 and flatten the
nested else. The check above it already guarantees two buffered bytes, and
both CRLF and "--" are two bytes, so it can never fire. Removing it is what
makes the new comment's claim readable straight off the code.
* Rename the timing test's locals to avoid a Windows macro
MSVC's <rpcndr.h>, pulled in by <windows.h>, defines `small` as `char`, so
`auto small = ...` failed to compile on the Windows jobs. Same class of
problem as the std::min / std::max collision.
* Add Server::CustomRoute() for HTTP methods outside the built-in set
parse_request_line validates the request method against a fixed whitelist and
rejects anything else with 400 before routing runs. That blocks WebDAV, where
PROPFIND, PROPPATCH, MKCOL, COPY, MOVE, LOCK and UNLOCK are ordinary methods
defined by RFC 4918, and it blocks extension methods such as UPnP's SUBSCRIBE.
The need has been open since #847.
Registering a handler is now what makes the server accept a method:
svr.CustomRoute("PROPFIND", "/dav/:id", handler);
Because custom methods go through the normal dispatch path, patterns work the
way they do for Get() and friends, and the request body is available in
req.body. Serving these methods through set_pre_routing_handler was never
enough: the body has not been read at that point, so PROPPATCH and LOCK, which
require one, could not be implemented at all.
A HandlerWithContentReader overload is available too. The content reader gate
in routing() also fires when a custom method carries no body, matching what
expect_content() does unconditionally for POST/PUT/PATCH/DELETE, so a body-less
PROPFIND (RFC 4918 treats one as allprop) reaches its handler instead of
falling through to 404.
Method names are validated as RFC 9110 tokens, and the ten built-in methods are
refused. Seven of them are dispatched by the if/else chain in routing() before
the custom tables are consulted, so a route registered for one could never
fire; CONNECT, TRACE and PRI carry protocol-level meaning this library does not
route. A refused registration makes is_valid() return false, so listen() fails
rather than starting a server holding a handler that would never run. This is
also why SSLServer::is_valid() now chains to Server::is_valid() instead of only
checking ctx_.
Servers that never call CustomRoute() keep the previous per-request cost: the
built-in method set is checked first and short-circuits, and the custom lookup
returns early on an empty map.
* Add cookbook recipe for custom HTTP methods
The CustomRoute() docs were a section inside S01, which pushed that page to 90
lines, the longest in the cookbook, and mixed a separate feature into a page
about registering GET/POST/PUT/DELETE handlers. Move the section into its own
recipe and give it room for the part that was missing: the OPTIONS handler
returning DAV: and Allow, which WebDAV clients probe for before anything else.
S01 goes back to 68 lines and keeps a pointer to the new page.
The recipe is titled after the API rather than after WebDAV, and says outright
that generating the 207 Multi-Status XML, interpreting Depth and managing locks
are the reader's job. Routing the method is all the library does.
S23 takes order 42, so the TLS, SSE and WebSocket recipes shift to 43-57. That
only moves the sort key. Filenames, the T01/E01/W01 labels, the published URLs
and every cross-reference are untouched.
Start-Process joins ArgumentList entries with spaces, so /DIR=C:\Program
Files\OpenSSL reached Inno Setup as /DIR=C:\Program and the install landed
there. Linking still succeeded, because the import libraries were present
under that path, and the failure surfaced only when gtest_discover_tests ran
the test binary: exit code 0xc0000135, DLL not found, since PATH pointed at
C:\Program Files\OpenSSL\bin.
Quote the value, and assert that the import libraries and runtime DLLs are
where we expect before exporting PATH, so a misplaced install fails loudly at
the install step instead of quietly at load time.
The Chocolatey openssl package hardcodes a versioned slproweb URL in its
install script, and slproweb keeps only the newest build of each OpenSSL
branch. Every OpenSSL release therefore deletes the file the current package
points at, and "windows with SSL" fails at the install step with a 404 until
someone respins the package. That is what broke the job today: the package is
still at 4.0.1 while slproweb has moved to 4.0.2.
slproweb publishes a JSON manifest of its current downloads, linked from the
download page and updated at the same time as the files themselves. Read that
and take the newest 64-bit 4.x installer from it, so the URL is always live.
The SHA512 in the manifest is verified before the installer runs.
The silent flags are the ones the Chocolatey package used. /DIR pins the
install location that the CMake step already finds, instead of relying on a
registry lookup. PATH and OPENSSL_CONF are exported the same way the package
set them.
Staying on 4.x is deliberate: it keeps this job on the OpenSSL 4.0 series
rather than dropping to the 3.6 that vcpkg would provide.
close() drained the peer's Close reply with its own frame read. If an
application reader thread was inside read() at that moment, two threads
parsed frames off one stream: read_websocket_frame()'s payload loop keeps
reading until it has the declared length, so bytes stolen by the drain
were silently replaced with bytes from further along the stream. The
in-flight message kept its correct length but got the wrong content.
Add a read_mutex_ that marks which thread owns the stream's read side.
read() holds it for the whole call. close() sends the Close frame, then
drains the peer's reply (RFC 6455 7.1.1) only if it can try_lock the
mutex; otherwise it returns immediately, leaving the stream entirely to
the thread already reading it. This also fixes close() blocking for the
full close timeout when a reader thread was parked waiting on a peer
that never replies.
Add WebSocketTest.CloseDoesNotStealBytesFromConcurrentRead, which drives
a raw TCP peer that stalls mid-payload to force the race; it fails
reliably against the old code and passes against the fix.
Update README-websocket.md: close() during a concurrent read() is now
supported.
A wss:// WebSocket enters a single TLS session from several threads: the
read path, the application's send()/close(), and the heartbeat ping thread.
The existing write_mutex_ only serializes writers, so a reader's SSL_read and
a writer's SSL_write (plus the SSL_peek in is_peer_closed() on the write path)
run concurrently on the same session. OpenSSL and the other backends forbid
concurrent access to one session, so this corrupts the record layer: messages
are silently dropped, and under ASan it shows up as a heap-buffer-overflow.
It affects wss:// only; plain ws:// is unaffected because the kernel allows
concurrent recv()/send() on a socket.
Route wss:// through a new WebSocketSSLStream that serializes every TLS call
with one per-stream mutex. The socket is kept non-blocking for the stream's
lifetime and each read()/write() performs a single non-blocking TLS call under
the lock, then waits for readiness with select() outside the lock. The lock is
therefore held only for CPU-bound work, so a reader blocked waiting for data
never stalls a concurrent sender.
Because the socket is non-blocking, a TLS call can stop needing either
direction, so read() also waits for writability on WantWrite and write() waits
for readability on WantRead. A read that shares its session with the send path
has to flush pending output before it can decrypt more input, and Mbed TLS
surfaces this on every mbedtls_ssl_read(). The read timeouts are atomic since
WebSocket::close() shortens them from the closing thread while the receive
thread is inside wait_readable().
SSLSocketStream is left untouched, so ordinary HTTP/HTTPS keeps its exact code
path and performance. The heartbeat ping thread also stays, so timer-driven
pings keep working as before.
Add test_websocket_thread_safety.cc, which drives send/close/heartbeat against
a concurrent reader over wss://. Built with ASan in CI, a regression surfaces
as a heap-buffer-overflow.
The harness had a single endpoint returning a 12-byte set_content() body.
That is the one case where the response line, the headers and the body
already share a single write(), so any change to the write path measured
as noise. Comparing a gather-write branch against its merge base reported
0.993x at p = 1.000 while the same branch moved static-file throughput by
a quarter and TLS throughput by nearly half in both directions.
The server now also serves a large set_content() body and small and large
files from a mount point, over HTTPS when a certificate is given, with
--path, --large-mib and --tls selecting the combination.
ab.sh now compiles the harness from the invoking worktree instead of each
ref's own copy, so both refs run an identical workload and a ref that
predates a harness change stays measurable. Only httplib.h varies, through
-I. --timeout is exposed because bombardier's 2s default aborts large TLS
responses, which then fails the non-2xx check.
Drop the now-redundant has_header guard (get_header_value already
returns "" for a missing header, which the length check rejects),
name the "Bearer " prefix once, and cite RFC 9110 to match the
file's convention. Convert the regression test to the table-driven
form used elsewhere in test.cc and move it out of the middle of
GetHeaderValueTest so that suite stays contiguous.
detail::parse_www_authenticate() assumed a single challenge starting at
the first space in the field value and read only its first occurrence,
so a Basic challenge listed before Digest (or split across two field
lines, as some servers do) hid the Digest challenge entirely, and a
second Digest challenge with different parameters (RFC 7616 offering
both SHA-256 and MD5) could mix params from both. Combine repeated
field lines the same way the other list-valued headers do, then split
on commas that aren't inside a quoted-string so a quoted realm can
contain a comma, and track which challenge each auth-param belongs to
by the auth-scheme token that starts it. Also require at least one
auth-param before reporting a Digest challenge as found, since an
empty challenge can't produce a usable Authorization header.
RFC 9110 7.8 defines Upgrade as a comma-separated list of protocols and asks
recipients to match each protocol-name case-insensitively; RFC 6455 4.2.1 asks
for a header field containing the value "websocket". Both handshake checks
instead read occurrence zero and required the whole field value to be exactly
"websocket", so a client offering "websocket, HTTP/3.0" -- or naming websocket
on a second Upgrade field line -- was answered 404 rather than 101.
This is the defect ffe2a1c fixed for Connection two lines below, and
has_header_token() is already called in both of these functions.
The client-side check loosens what we accept back from a server, which is the
same reading: a server answering 101 may name websocket alongside another
protocol, and rejecting that handshake was ours to get wrong.
The four ExpectTokenTest cases landed inside the #ifndef _WIN32 that guards the
10 GiB content-provider test below them, so Windows never compiled them and the
green Windows jobs said nothing about the fix. Nothing in them is POSIX-only --
they use the same helpers as ConnectionTokenTest, which sits outside any guard
-- so move them above the guard.
probe_expect() re-implemented send_request(), down to the create_client_socket
argument list. Call send_request() instead, with Connection: close so its read
loop ends at the response rather than idling to the read timeout; the Connection
check runs before the Expect block, so it does not disturb what is under test.
Move the new Content-Encoding case below its siblings. Appending it to the tail
of the comment block left the "whole token" paragraph reading as documentation
for a test about repeated field lines. The paragraph above it had been detached
from KnownEncodingWithoutSupportIsReported the same way one commit earlier; put
that one back too.
RFC 9110 Section 10.1.1 defines Expect as a comma-separated list, states that
its value is case-insensitive, and requires a server that receives a
100-continue expectation in an HTTP/1.0 request to ignore it. Comparing the
whole field value against "100-continue" met none of those.
An HTTP/1.0 request asking for 100-continue was answered with a 100 (Continue)
interim response, which that section forbids. "100-Continue" and
"100-continue, foo" were both read as no expectation at all, so a client that
waits for the interim response before sending its content waited for a response
that was never coming.
Route the check through has_header_token(), which walks every field line and
compares complete tokens case-insensitively, and skip it for HTTP/1.0. An
expectation cpp-httplib does not recognize is still ignored rather than
refused; the 417 the section offers for one is a MAY, not a requirement.
RFC 9110 Section 5.3 makes a Content-Encoding spread over several field lines
the same message as the comma-joined one, so the two have to be read the same
way. Reading occurrence zero did not: a response carrying "gzip" on two field
lines was decoded as a single gzip coding, so a body the sender says was
encoded twice came back after one pass -- still compressed, but presented to
the caller as decoded. The same value written as "gzip, gzip" on one line took
the pass-through path instead.
Read the combined value at both sites. A value naming several codings matches
none of the ones cpp-httplib implements, so both representations now take the
pass-through path that prepare_content_receiver() already documents for an
unrecognized coding.
This does mean a sender that repeats "Content-Encoding: gzip" on two lines for
a body it gzipped once no longer has that body decoded. There is no way to tell
that sender apart from one that really did encode twice, and the conservative
reading is the one the field value states.
is_brotli_encoding() and is_zstd_encoding() searched the Content-Encoding
value for "br" and "zstd" as substrings, while is_zlib_encoding() beside them
compared the whole value. So "fibre" and "librarian" were read as Brotli and
"x-zstd-ish" as Zstandard, and "gzip, br" -- a value naming two codings, which
cpp-httplib does not support -- was labeled Brotli and run through a Brotli
decompressor over gzip data.
RFC 9110 8.4.1 defines a content coding as a token, so compare the whole value
case-insensitively as the zlib check already does. A value naming several
codings no longer matches any of them and takes the pass-through path
prepare_content_receiver() already documents for an unrecognized coding.
contains_case_ignore() has no callers left.
RFC 9110 Section 7.6.1 defines Connection as a comma-separated list of
case-insensitive connection options, and Section 5.3 lets that list be split
across several field lines. Comparing the whole field value against a single
option gets both wrong.
A client sending "Connection: keep-alive, close" was answered without a
Connection header and its socket was kept open, so the close it asked for was
never performed and never announced. An HTTP/1.0 client asking for keep-alive
only got it by spelling the option exactly "Keep-Alive"; the lowercase form
everyone actually sends closed the connection instead.
Route the five Connection checks through has_header_token(), which already
walks every field line and compares complete tokens. Expect is left alone:
matching "100-continue" as a token would make an unrecognized expectation
alongside it look acceptable, where Section 10.1.1 asks for 417.
It was defined as a static inline above the border line so that the split
build would not turn it into an exported symbol of the shared library, which
kept abidiff from reporting an added function. That put an internal helper's
location at the mercy of a CI check rather than of where it belongs: it reads
a header field the same way get_header_value() and get_combined_header_value()
do, and it is the closest sibling of the latter, both being about a list-valued
field spread over several field lines.
Define it as a plain inline beside them and forward-declare it with the
split() family it calls. Adding a symbol is a source and binary compatible
change, so let abidiff report it.
RFC 9110 Section 5.2 and 5.3 define the combined value of repeated field
lines as their values joined by commas in the order they were received.
Several call sites read only the first occurrence and then split that on
commas, so whatever the later field lines carried was silently dropped: an
acceptable media type or content coding, an ETag, a WebSocket subprotocol, a
declared trailer name, or an address a proxy appended as its own line rather
than by extending the one it received.
Add detail::get_combined_header_value() and use it for Accept,
Accept-Encoding, If-None-Match, Sec-WebSocket-Protocol, Trailer and
X-Forwarded-For. Empty field lines are skipped so the combined value never
starts with a bare comma, which parse_accept_header() rejects outright.
Also drop the now-dead manual trimming in parse_trailers() and replace the
istringstream-based subprotocol tokenizer with detail::split(); split()
already trims each token and skips empty ones.
* Match the Connection "Upgrade" token exactly in WebSocket handshakes
The server and the client both looked for "upgrade" as a substring of the
Connection field value, so "notupgrade", "upgrade-not" and "xupgrade" all
passed as the standalone token the handshake requires. RFC 6455 4.2.1 asks
for an ASCII case-insensitive token match, and a value split across several
Connection lines was missed entirely because only the first line was read.
Parse the field as the comma-separated token list it is, across every line,
and reuse the same helper for the server request check and the client
response check.
Reported by gb1dev.
* Tidy up the Connection token helper
Move has_header_token() out of the WebSocket-only detail block and next to
the other header field helpers, forward-declaring it beside split(). Use the
existing split_find(), which drops the manual found flag and stops at the
first matching token.
Drive the client-side test from an ordinary Server route answering 101 with
a bad Connection value, rather than the hand-rolled listening socket copied
from the test above it.
* Keep the Connection token helper out of the split build's ABI
The split build strips inline from everything below the border line, so a
helper defined there becomes an exported symbol of the shared library and
abidiff reports it as an added function. The tests also could not see
is_websocket_upgrade() or websocket_accept_key(), since neither is declared
in the part of the header that survives the split.
Define has_header_token() as a static inline above the border, next to the
split() declarations its two call sites already sit below, and declare the
two WebSocket helpers the way ws::impl::read_websocket_frame() already is.
The shared library's exported symbols are now identical to master's.
decode_uri was a byte-for-byte copy of decode_uri_component: it decoded every
%XX, including escapes of the reserved characters that encode_uri leaves
literal. So decode_uri was not the inverse of encode_uri and promoted an
escaped delimiter into a real one -- decode_uri("http://h/a%2Fb") returned
"http://h/a/b". Keep escapes of the reserved set encode_uri preserves, matching
JS decodeURI; non-reserved escapes still decode.
Server::make_matcher() built a std::regex for every pattern that did not
contain "/:", even though most route patterns are plain literals with no
regular expression syntax in them. Matching those went through
std::regex_match on every request, for every registered route the
dispatcher scanned before reaching the one that matches.
PathParamsMatcher already performs an exact literal comparison when it
captures no parameter, so no new matcher class is needed: a pattern with
no regex metacharacter can simply use it. Add an early return for the
zero parameter case in PathParamsMatcher::match(), and select the matcher
by also looking for the 14 ECMAScript metacharacters instead of only for
"/:". Path params keep taking precedence, so a pattern that mixes both,
such as "/users/:id/(.*)", is unaffected.
Measured with clang -O2 on macOS, scanning routes that all miss until the
last one: at 100 routes a scan drops from 10.3us to 0.46us, and end to end
throughput rises by about 24%. At 1000 routes throughput is roughly 3
times higher. Registering 5000 routes drops from about 3.0ms to about
0.9ms, since no std::regex is built for literal patterns.
This also keeps CPPHTTPLIB_REGEX_ROUTE_PATH_MAX_LENGTH confined to the
routes it is meant for. That limit rejects overlong paths before calling
std::regex_match, but until now every literal route was a RegexMatcher
too, so a literal route longer than the limit stopped matching even though
no regular expression was involved. Literal routes no longer go through
RegexMatcher, so only real regex routes are capped.
Patterns containing a metacharacter keep their current behavior, so
"/index.html" still matches "/indexXhtml" the way it always has. One
visible change: a literal route no longer populates Request::matches,
which is now a default constructed std::smatch. Path parameter routes
have always behaved that way, and Request::matches only carries useful
information for regex routes.
RegexMatcher::match() called std::regex_match() directly on the
attacker-controlled request path. For quantified patterns such as "(.*)",
std::regex_match's recursive backtracking implementation (most acute on
libstdc++) recurses roughly once per matched character, so a long enough
path can exhaust the calling thread's stack and crash the process. Verified
against real GNU libstdc++: under the default thread stack size, a path of
a couple thousand characters against a simple quantified route pattern
reliably crashed the process, well within the existing 8192-byte request
URI limit.
Add CPPHTTPLIB_REGEX_ROUTE_PATH_MAX_LENGTH (default 256) and reject paths
longer than it before ever calling std::regex_match, treating them as a
non-match instead. Confirmed the fix eliminates the crash under the same
libstdc++ build and default stack size that reproduced it.
For a non-SSL request with neither Content-Length nor Transfer-Encoding,
Server::read_content_core() fell back to reading raw wire bytes with
detail::read_content_without_length() directly, bypassing the decompressor
wrapper that the length-framed and chunked paths already use. As a result,
payload_max_length only bounded the compressed bytes read off the socket,
not the decompressed size a handler could produce from them.
Route this fallback through detail::read_content(..., decompress=true)
instead, the same helper already used below for the length-framed and
chunked cases, so the decompressed-size guard applies uniformly.
* Gracefully drain socket before close in Server::process_and_close_socket
Closing a connection while the receive queue still has unread data,
or while bytes are still in flight, can make the OS send an abortive
RST instead of a graceful FIN. On Windows this surfaces as
WSAECONNABORTED/WSAECONNRESET on the peer's read, which can make an
otherwise fully-written response look like a failed request -- a
likely contributor to the ServerTest.HTTP2Magic flakiness tracked in
#2533.
Add detail::close_socket_gracefully(), which half-closes the write
side, drains any queued/in-flight bytes (bounded to 100ms / 1MB),
then performs the final shutdown+close. Use it in
Server::process_and_close_socket.
Root cause and fix mechanism identified by @Hyukya in #2533.
* Rename close_socket_gracefully to drain_and_close_socket
'gracefully' already means something specific in this codebase: whether
to send a TLS close_notify before closing (shutdown_ssl's
shutdown_gracefully param, ClientImpl::disconnect(gracefully),
tls::shutdown(session, graceful)). Reusing the word for an unrelated
TCP-level drain-before-close made the new function read as part of that
TLS machinery when it isn't. Rename it to describe what it does instead,
matching the existing close_socket/shutdown_socket and
WebSocketClient::shutdown_and_close naming.
The stricter ws::Result error checks added in 6018c7f and 86d0210
exposed two backend-parity bugs in setup_client_tls_session(), shared
by SSLClient and WebSocketClient since their TLS setup was merged:
- enable_server_hostname_verification(false) had no effect on Mbed TLS
or wolfSSL for DNS hosts: mbedtls_ssl_set_hostname() and
wolfSSL_check_domain_name() bind SNI and handshake-time identity
checking together, so the identity check ran regardless of the
option, failing the handshake before the post-handshake
server_hostname_verification check was ever reached.
- On a genuine wrong-hostname failure, Mbed TLS reported the generic
Error::SSLServerVerification instead of
Error::SSLServerHostnameVerification, because
MBEDTLS_ERR_X509_CERT_VERIFY_FAILED was mapped without looking at
which verify flag actually caused it.
Fixes:
- set_sni() now takes a verify_hostname flag. wolfSSL skips
wolfSSL_check_domain_name() when it's false. Mbed TLS can't request
SNI without also arming the CN/SAN check, so it installs a verify
callback that masks the mismatch flag instead - a self-contained one
when the session has no user verify callback of its own, so it never
reads the process-wide set_verify_callback() slot another client may
have populated (this was caught by ASAN as a stack-use-after-scope:
VerifyCallbackTest.VerifyContextFields leaves a dangling lambda
there because MbedTlsSession never had a reason to consult it
before).
- map_mbedtls_error() now takes the handshake's verify flags and
reports HostnameMismatch when CN/SAN mismatch is the only one set,
matching the wolfSSL mapping and the post-handshake identity check.
- The duplicated verify-flags/error-mapping/backend_code logic in
connect() and connect_nonblocking() is factored into
fill_mbedtls_tls_error(); the duplicated flag-clearing in the two
verify callbacks is factored into mbedtls_clear_cn_mismatch(); both
use the existing hostname_mismatch_code() accessor instead of the
raw Mbed TLS macro.
Also tightens SSLClientTest.ServerHostnameVerificationError_Online to
assert the specific error code now that all three backends agree,
rather than accepting Mbed TLS's old fallback value.
Verified full non-online suite green on OpenSSL (791), Mbed TLS (737),
and wolfSSL (735), plus the split build, plus the Online
hostname-mismatch test against badssl.com on all three backends.
WebSocketClient's TLS setup already threaded
ClientTlsSessionOptions::server_hostname_verification through
setup_client_tls_session(), the same path SSLClient uses, but never
exposed a way to set it: create_stream() called setup_client_tls_session()
without an options argument, so the default (verification on) was the
only reachable value.
Add the public setter, mirroring ClientImpl/SSLClient/Client, and wire
it into create_stream()'s ClientTlsSessionOptions. Last open item from
issue #2531's WebSocketClient/SSLClient API alignment.
Issue #2531 asked for connect() to expose error detail the way
ClientImpl/SSLClient do via Result, instead of collapsing every failure
into a bare bool. The groundwork (detail::ClientTlsSessionError) was
already laid during the WebSocketClient/SSLClient dedup but left
unwired.
- Add httplib::ws::Result: explicit operator bool(), error(), and
flattened upgrade-response accessors (status(), headers(),
get_header_value(), has_header()); ssl_error()/ssl_backend_error() on
SSL builds.
- Add Error::WebSocketHandshake for upgrade-validation failures
(non-101 status, bad Sec-WebSocket-Accept, bad Upgrade/Connection
headers).
- Extract detail::parse_status_line from ClientImpl::read_response_line
and reuse it in read_websocket_upgrade_response, replacing the
previous "HTTP/1.1 101" substring match with a proper parse. Non-101
responses now surface their status and headers instead of being read
and discarded.
- Wire WebSocketClient::create_stream() to capture ClientTlsSessionError
so TLS failures (SSLServerVerification, SSLServerHostnameVerification,
...) reach the caller with backend error codes.
- Update tests and README-websocket.md accordingly.
This is a source-breaking change for callers that assign the result to
bool (e.g. bool ok = cli.connect();); if (cli.connect()) and gtest's
ASSERT_TRUE/EXPECT_FALSE(...) macros are unaffected since operator bool
still participates in contextual conversion.
T04 (mTLS) had grown a "WebSocketClient" subsection describing
wss:// client certificates, and c12/t02 were getting similar
WebSocketClient asides for timeouts and CA paths. The Cookbook's
own index already separates WebSocket into its own category
(W01-W04) from TLS/Security (T01-T05) and Client (C01-C19), so
burying WebSocketClient specifics inside those pages fought the
site's structure.
Move that content into two new recipes under the WebSocket
category instead:
- W05: wss:// TLS setup (set_ca_cert_path CA directory parity,
PemMemory client certificate)
- W06: WebSocketClient's three timeouts, including the recently
added chrono overloads
T04, T02, C12, and W01 now carry a single reference link to the
new pages instead of duplicated explanations, matching the site's
existing cross-link convention.
While rewriting T04's client-side section, noticed it documented
SSLClient's file-path constructor but not its PemMemory one, even
though the server-side section covered both forms for SSLServer.
Added the missing PemMemory example so both sides are symmetric.
WebSocketClient::set_connection_timeout (both the time_t and
chrono overloads) was missing from README-websocket.md's API
reference and the timeout example, even though set_read_timeout
and set_write_timeout were both listed.
README.md never documented the PemMemory in-memory constructor
that SSLServer and SSLClient both have, so mTLS setup only showed
the file-path form. Add a "Mutual TLS (mTLS)" section covering
both forms for server and client, and note that
ws::WebSocketClient's wss:// constructor takes the same PemMemory
struct.
Also note, next to Client::set_interface, that WebSocketClient has
the same method, matching the existing cross-reference for
set_hostname_addr_map right below it.
Adds ws::WebSocketClient::PemMemory and a constructor overload that
installs an in-memory client certificate on the TLS context, enabling
mutual TLS for wss:// connections. The certificate is silently ignored
for ws:// URLs, consistent with the existing TLS-only setters such as
set_ca_cert_path().
Part of the interface alignment discussed in #2531.
Temporary instrumentation to track the intermittent windows-without-SSL
failures reported in #2533. When the job fails on a push, it posts the
run URL, commit, and per-shard failed-test lines as a comment on the
issue, building up failure-pattern history automatically.
This should be removed once the root cause is found and fixed.
SSLClient::initialize_ssl kept its own copy of the session setup that
detail::setup_client_tls_session already implemented for WebSocketClient.
Extend the shared function with the pieces only SSLClient needed - a session
verifier, an independent hostname verification flag, the context mutex,
Windows Schannel verification and error details - and let initialize_ssl
build a ClientTlsSessionOptions and call it. All of them default, so
WebSocketClient's call site is unchanged.
This settles one difference between the two: WebSocketClient used to call
tls::set_hostname for named hosts, which on OpenSSL turns on verification
during the handshake, while SSLClient always set SNI only and verified
post-handshake. The shared function now does the latter for both, so
tls::set_hostname loses its last caller and goes away, as does the
write-only SSLClient::verify_result_.
Certificate verification with a host name rather than an IP literal was the
one combination the WebSocket tests never covered, and it is exactly the
path this normalizes. WebSocketSSLDnsHostTest fills that in; cert2 gains a
DNS:localhost SAN so a name can be verified against it.
ClientImpl::prepare_default_headers and WebSocketClient::prepare_default_headers
each built the Host value with the same AF_UNIX special case and appended the
same User-Agent. Move both into detail:: so there is one copy.
The Host helper returns only the value, because the two callers disagree on
where it goes: ClientImpl prepends it per RFC 9110 5.3, WebSocketClient appends.
The User-Agent helper takes the Request, since it has to consult and set a
header rather than compute a string, and it stays inside ClientImpl's
content_receiver branch so that path keeps sending no User-Agent.
WebSocketClient::set_ca_cert_path took a single path and create_stream()
hardcoded an empty directory when calling detail::load_client_ca_config, while
ClientImpl has always accepted (ca_cert_file_path, ca_cert_dir_path = ""). Give
WebSocketClient the same signature and store the directory, so both clients
configure CA loading identically. The one-argument form is unchanged for
callers.
Also note at both call sites why the "load the CA config once" guard differs:
SSLClient needs call_once because one client serves concurrent requests, and
WebSocketClient does not because connect() is not safe to call concurrently
anyway.
WebSocketClient only accepted timeouts as (time_t sec, time_t usec), while
ClientImpl has taken std::chrono::duration overloads for its read, write and
connection timeouts for a long time. Add the same three overloads, forwarding
through the existing detail::duration_to_sec_and_usec helper so the split
matches ClientImpl exactly.
The template bodies go above the first split.py BORDER, next to the class, so
that the .h/.cc split keeps them in the header where instantiation needs them.
Timeout=3,// Read timeout elapsed; connection still open
};
```
Returned by `read()`. Since `Fail` is `0`, the result works naturally in boolean contexts — `while (ws.read(msg))` continues until the connection closes. When you need to distinguish text from binary, check the return value directly.
`Timeout` is only returned for a read timeout you set yourself with `set_read_timeout()`. It means the timeout elapsed on a message boundary: nothing was consumed and the connection is still open, so you can send on it and read again. The compile-time defaults (`CPPHTTPLIB_WEBSOCKET_SERVER_READ_TIMEOUT_SECOND`, 300 seconds on the server; a client waits forever) are a backstop against a peer that has gone quiet, not a request for control: when one of them elapses, `read()` returns `Fail` and closes the connection, so code that never calls `set_read_timeout()` can keep using `while (ws.read(msg))`.
**`msg` is left untouched on `Timeout`.** Because `Timeout` is non-zero, `while (ws.read(msg))` keeps looping — with the *previous* message still in `msg`. Once you set a read timeout, test the result instead:
The check above runs after the handshake, so the client sees a successful upgrade followed by a close frame. To refuse the upgrade itself with an HTTP status, use a pre-routing or pre-request handler. Both run before the `101 Switching Protocols` response, and `req.matched_route` is available in the pre-request handler:
| `CPPHTTPLIB_WEBSOCKET_MAX_MISSED_PONGS` | `0` (disabled) | Close the connection after N consecutive unacked pings |
@@ -390,7 +475,7 @@ The server side has the same `set_websocket_max_missed_pongs()`.
With the default ping interval of 30 seconds, `max_missed_pongs = 2` detects a dead peer within ~60 seconds. The counter is reset every time a Pong frame is received, so the mechanism only works when your code is actively calling `read()` — exactly the pattern a normal WebSocket client already uses.
**The default is `0`**, which means "never close the connection because of missing pongs." Pings are still sent on the heartbeat interval, but their responses are not checked. Even so, a dead connection does not linger forever: while your code is inside `read()`, `CPPHTTPLIB_WEBSOCKET_READ_TIMEOUT_SECOND` (default **300 seconds = 5 minutes**) acts as a backstop and `read()` fails if no frame arrives in time. `max_missed_pongs` is the knob for detecting an unresponsive peer faster than that 5-minute fallback.
**The default is `0`**, which means "never close the connection because of missing pongs." Pings are still sent on the heartbeat interval, but their responses are not checked. On the server side a dead connection still does not linger: while a handler is inside `read()`, `CPPHTTPLIB_WEBSOCKET_SERVER_READ_TIMEOUT_SECOND` (default **300 seconds = 5 minutes**) acts as a backstop. A client has no such backstop — it waits forever unless you set a read timeout — so there `max_missed_pongs` is what notices an unresponsive peer at all. On either side it is also the knob for noticing one *faster* than the 5-minute fallback.
## Threading Model
@@ -410,6 +495,14 @@ svr.new_task_queue = [] {
Choose sizes that account for both your expected HTTP load and the maximum number of simultaneous WebSocket connections.
### Calling from Multiple Threads
A single `WebSocket` (server-side) or `WebSocketClient` handle is shared by three potential callers: the thread running your handler (or holding the client), the heartbeat thread, and, if your code does its own thing, a separate thread calling `send()`/`close()` while another thread is blocked in `read()`.
**Supported**: calling `read()` from one thread while calling `send()`/`close()` from another. This is the common pattern for a client that reads incoming messages in a loop on one thread and sends from elsewhere (e.g. a UI thread). A message that is in flight when `close()` is called still arrives intact; `close()` sends the Close frame and returns, leaving the connection's read side to the thread that owns it, so it does not block waiting for the peer's Close reply in that case. The heartbeat thread's automatic pings use the same `send()` path internally, so they are safe to run concurrently with your `read()` loop too — for `wss://` this requires every TLS call on a connection to be serialized internally, which cpp-httplib does for you.
**Not supported**: calling `read()` from two threads at the same time on the same handle. The calls are serialized rather than left to corrupt each other, but which thread receives which message is unspecified, so there is nothing useful to build on it.
## Protocol
The implementation follows [RFC 6455](https://tools.ietf.org/html/rfc6455):
Both `SSLServer` and `SSLClient` also accept an in-memory `PemMemory` struct instead of file paths — handy when certs come from an environment variable or a secrets manager:
`httplib::ws::WebSocketClient` has the same `PemMemory` constructor for `wss://` connections. See [README-websocket.md](README-websocket.md) for details.
### Peer Certificate Inspection
On the server side, you can inspect the client's peer certificate from a request handler:
@@ -272,6 +307,39 @@ int main(void)
`Post`, `Put`, `Patch`, `Delete` and `Options` methods are also supported.
### Custom HTTP methods
Methods outside the built-in set are rejected with `400 Bad Request` unless a handler is registered for them with `CustomRoute`. This covers the WebDAV methods of RFC 4918, `SUBSCRIBE` and friends from UPnP, and any other extension method.
Patterns work exactly as they do for `Get` and the other methods, so regular expressions and path parameters are both available.
Note the following:
* The method name must be a valid HTTP method token (RFC 9110) and must be registered before `listen()` is called.
* `GET`, `HEAD`, `POST`, `PUT`, `DELETE`, `CONNECT`, `OPTIONS`, `TRACE`, `PATCH` and `PRI` cannot be registered this way. Use the dedicated methods above instead.
* A rejected registration makes `is_valid()` return `false`, and `listen()` then fails rather than starting a server with a route that would never fire.
* Static file serving and WebSocket upgrades remain `GET`/`HEAD` only.
* `Allow` and the WebDAV `DAV:` header are not generated automatically. Register an `Options` handler if clients need them.
### Bind a socket to multiple interfaces and any available port
```cpp
@@ -279,6 +347,25 @@ int port = svr.bind_to_any_port("0.0.0.0");
svr.listen_after_bind();
```
### Port sharing and exclusive binding
By default, the server socket enables address/port reuse: `SO_REUSEPORT` where it is available (Linux, macOS), and `SO_REUSEADDR` otherwise (Windows). A restarted server can bind again immediately, but binding to a port that another server is already listening on also succeeds, and connections are distributed between them.
If you want `listen()` to fail when the port is already in use, replace the default socket options with `set_socket_options`:
> Setting only `SO_REUSEADDR` is not enough on Windows. There, `SO_REUSEADDR` allows two sockets that both set it to bind to the same port, so use `SO_EXCLUSIVEADDRUSE` instead.
The pre-compression logger is only called when compression would be applied. For responses without compression, only the access logger is called.
For a static file response (see [Static file compression](#static-file-compression)), `res.body` is empty when the logger runs. The bytes are still on disk at that point, not in memory.
#### Error Logging
Error loggers capture failed requests and connection issues. Unlike access loggers, error loggers only receive the Error and Request information, as errors typically occur before a meaningful Response can be generated.
├─ expect_100_continue_handler (when the request has "Expect: 100-continue")
│
├─ route matching → req.matched_route is set
│
├─ pre_request_handler route matched, body NOT read yet
@@ -498,6 +588,10 @@ Request received
Use `pre_routing_handler` to reject a request as early as possible, before the route is known. Use `pre_request_handler` for route-specific checks, since `req.matched_route` is available and the body has not been read yet.
For a request with `Expect: 100-continue`, the `100 Continue` response is not sent until the body is about to be read. A request rejected before that point (by `pre_routing_handler`, `pre_request_handler`, or because no route matched) gets its final response without `100 Continue`, so the client never sends the body.
A WebSocket upgrade request that matches a route registered with `svr.WebSocket()` takes a shorter path: `pre_routing_handler`, then route matching (`req.matched_route` is set), then `pre_request_handler`, then the WebSocket handler. If either hook returns `Handled`, its response is sent as a regular HTTP response and the connection is not upgraded. Once the connection is upgraded, `post_routing_handler` does not run.
### Response user data
`res.user_data` is a type-safe key-value store that lets pre-routing or pre-request handlers pass arbitrary data to route handlers.
By default, the server sends a `100 Continue` response for an `Expect: 100-continue`header.
By default, the server accepts an `Expect: 100-continue` header and sends a `100 Continue` response when it starts reading the request body. If the request is answered without reading the body, `100 Continue`is not sent and the connection is closed after the response.
The handler runs before `pre_routing_handler`. Returning `100` lets the request proceed; returning any other status sends that status as the final response and closes the connection.
```cpp
// Send a '417 Expectation Failed' response.
@@ -1290,6 +1392,8 @@ res->status; // 200
cli.set_interface("eth0"); // Interface name, IP address or host name
```
The same method is available on `httplib::ws::WebSocketClient`.
### Override the connection target for a hostname
`set_hostname_addr_map` redirects where the socket connects, without changing
@@ -1357,6 +1461,29 @@ httplib::Server svr;
svr.listen("127.0.0.1", 8080);
```
## Ordered Headers, Query Parameters, and Form Data
`Headers`, `Params`, `FormFields`, and `FormFiles` preserve the order entries were received (for a parsed request) or inserted (for one you build yourself). Earlier versions stored these in `std::multimap` or `std::unordered_multimap`, which either sorted entries by key or gave no ordering guarantee at all for repeated keys. RFC 9110 §5.3 and RFC 7578 §5.2 both require the original order to be preserved, so this is now guaranteed rather than incidental.
```c++
// A request with two Accept-Encoding lines...
// Accept-Encoding: gzip
// Accept-Encoding: br
// ...visits "gzip" before "br", not the other way around.
for (auto it = req.headers.equal_range("Accept-Encoding").first;
it != req.headers.end(); ++it) {
std::cout << it->second << std::endl;
}
// get_header_value(key, id) reaches a specific one directly.
auto second = req.get_header_value("Accept-Encoding", 1); // "br"
```
`Headers` matches field names case-insensitively, as before. `Params`, `FormFields`, and `FormFiles` are case-sensitive.
> [!NOTE]
> Iterators on these containers follow `std::vector` rules: inserting a new entry invalidates existing iterators. Code that keeps an iterator across a call to `insert()`/`emplace()` needs to re-fetch it afterward.
## Payload Limit
The maximum payload body size is limited to 100MB by default for both server and client. You can change it with `set_payload_max_length()` or by defining `CPPHTTPLIB_PAYLOAD_MAX_LENGTH` at compile time. Setting it to `0` disables the limit entirely.
@@ -1373,6 +1500,45 @@ The server can apply compression to the following MIME type contents:
- application/protobuf
- application/xhtml+xml
A response that already carries `Content-Encoding` is sent as it is. A handler serving content it encoded itself, an asset compressed at build time for instance, keeps its own coding and its own bytes:
This holds for every kind of response, including the file-backed ones below.
`Vary: Accept-Encoding` is added only to responses the server encoded itself. A handler that chooses between an encoded and an identity representation by reading `Accept-Encoding` should set the field itself, so that shared caches keep the two apart.
### Static file compression
Responses served from a file, whether through `set_mount_point()` or `Response::set_file_content()`, are sent as is by default. Turn compression on for them with:
```c++
svr.set_static_file_compression(true);
```
Only files within a size range are compressed, and both ends of it can be moved:
The lower bound defaults to 1400 bytes. A response that already fits in a single 1500-byte MTU is not delivered any faster for being smaller, and a file of a few bytes comes back larger than it went in, since gzip's header and trailer outweigh what deflate saves. `0` compresses everything down to a single byte, and `CPPHTTPLIB_STATIC_FILE_COMPRESSION_MIN_LENGTH` sets the default at compile time. An empty file is never compressed regardless.
The upper bound defaults to 4MB, and exists for a different reason: the file is compressed per request, and the compressed bytes are held in memory until the response has been written, so the peak cost scales with the number of requests in flight. It is a bound on what one request can cost, not a statement about how well large files compress, which is why raising it is reasonable when the files are known and the traffic is not. `0` removes the limit, and `CPPHTTPLIB_STATIC_FILE_COMPRESSION_MAX_LENGTH` sets the default at compile time.
A compressed response keeps its `Content-Length`, so `HEAD` still reports the size a `GET` would return. Two details are worth knowing:
- Range requests are answered from the uncompressed representation, so `Content-Range` keeps naming the file's own bytes.
- The `ETag` carries the coding it belongs to (`W/"...-gzip"`), so a client that cached the compressed form revalidates against the right validator.
Content providers registered with `set_content_provider()` are not covered. Feeding one through a compressor would hold each write back until the compressor's window filled, which breaks providers that produce their body incrementally. Use `set_chunked_content_provider()` to compress a generated body.
### Zlib Support
'gzip' compression is available with `CPPHTTPLIB_ZLIB_SUPPORT`. `libz` should be linked.
@@ -1421,7 +1587,6 @@ res->body; // Compressed data
Unix Domain Socket Support
--------------------------
Unix Domain Socket support is available on Linux and macOS.
> **Warning:** The read timeout covers a single receive call — not the whole request. If data keeps trickling in during a large download, the request can take half an hour without ever hitting the timeout. To cap the total request time, use [C13. Set an overall timeout](../c13-max-timeout).
> For WebSocket client timeouts, see [W06. Set Timeouts](../w06-websocket-timeouts).
With `httplib::Server`, you register a handler per HTTP method. Just pass a pattern and a lambda to `Get()`, `Post()`, `Put()`, or `Delete()`.
With `httplib::Server`, you register a handler per HTTP method. Just pass a pattern and a lambda to `Get()`, `Post()`, `Put()`, or `Delete()`. For methods outside the built-in set, such as WebDAV's `PROPFIND`, use `CustomRoute()`.
## Basic usage
@@ -64,3 +64,5 @@ To add a response header, use `res.set_header("Name", "Value")`.
> **Note:**`listen()` is a blocking call. To run it on a different thread, wrap it in `std::thread`. If you need non-blocking startup, see [S18. Control startup order with `listen_after_bind`](../s18-listen-after-bind).
> To use path parameters like `/users/:id`, see [S03. Use path parameters](../s03-path-params).
> For methods outside the built-in set, such as WebDAV's `PROPFIND`, see [S23. Handle custom HTTP methods](../s23-custom-methods).
Only a small chunk sits in memory at any moment, so gigabyte-scale files are no problem.
## Count the parts yourself
There is a cap on the number of parts, `CPPHTTPLIB_MULTIPART_FORM_DATA_FILE_MAX_COUNT` (1024 by default), but it only applies to the buffered path, where every part is accumulated into `req.form`. The `ContentReader` keeps nothing on the library side, so the cap does not apply here.
If you want an upper bound, count the parts yourself and return `false` from the header callback. The parser stops right there.
When `content_reader` returns `false`, set the response status yourself. The rest of the body is left unread and the connection is closed, so a client that is still sending sees the connection drop.
> **Warning:** When you use `HandlerWithContentReader`, `req.body` stays **empty**. Handle the body yourself inside the callbacks.
> For the client side of multipart uploads, see [C07. Upload a file as multipart form data](../c07-multipart-upload).
> **Note:** Tiny responses barely benefit from compression and just waste CPU time. cpp-httplib skips compression for bodies that are too small to bother with.
## Static files need to be opted in
Files served as they are, through `set_mount_point()` or `Response::set_file_content()`, are not compressed by default. Turn it on with:
```cpp
svr.set_static_file_compression(true);
```
Only files within a size range are compressed, and both ends of it can be moved:
The lower bound defaults to 1400 bytes. A response that already fits in a single 1500-byte MTU is not delivered any faster for being smaller, and a file of a few bytes comes back larger than it went in, because gzip's header and trailer outweigh what deflate saves.
The upper bound defaults to 4MB and exists for a different reason: the file is compressed on every request, and the compressed bytes stay in memory until the response has been written, so the peak cost scales with the number of requests in flight. It bounds what a single request can cost, and says nothing about how well large files compress, so raising it is reasonable when the files are known and the traffic is not.
Either bound takes `0` to turn it off, and each has a compile-time default (`CPPHTTPLIB_STATIC_FILE_COMPRESSION_MIN_LENGTH`, `CPPHTTPLIB_STATIC_FILE_COMPRESSION_MAX_LENGTH`).
A compressed response keeps its `Content-Length`, so `HEAD` reports the same size a `GET` would. Two details to know: Range requests are answered from the uncompressed representation, and the `ETag` carries the coding it belongs to, as in `W/"...-gzip"`.
Content providers registered with `set_content_provider()` are not covered. Running one through a compressor holds each write back until the internal buffer fills, which stalls providers that build their body a piece at a time. To compress a generated body, use `set_chunked_content_provider()`.
> **Note:** The size range covers static files only. A body passed to `set_content()` is compressed whenever the client accepts it and the MIME type is compressible, however small it is, so a response of a few bytes ends up larger than it started. Decide in the handler if you want to avoid that.
> For the client-side counterpart, see [C15. Enable compression](../c15-compression).
`matched_route` is the pattern **before** path parameters are expanded (e.g. `/admin/users/:id`). You compare against the route definition, not the actual request path, so IDs or names don't throw you off.
The pre-request handler also runs for routes registered with `svr.WebSocket()`. It is called before the `101 Switching Protocols` response, so returning `Handled` sends your HTTP response (such as a 403) and the connection is never upgraded.
`bind_to_port()` returns `false` on failure — typically when the port is already taken. Always check it.
`bind_to_port()` returns `false` on failure, for example when you don't have permission to bind to the port. Always check it.
```cpp
if (!svr.bind_to_port("0.0.0.0", 8080)) {
std::cerr << "port already in use" << std::endl;
std::cerr << "bind failed" << std::endl;
return 1;
}
```
`listen_after_bind()` blocks until the server stops and returns `true` on a clean shutdown.
## Detect a port that's already in use
With the default settings, you can actually bind to a port another server is already using. That's because cpp-httplib sets `SO_REUSEPORT` (Linux, macOS) or `SO_REUSEADDR` (Windows) on the server socket. A restarted server can bind again right away. The flip side is that a second server on the same port starts without an error, and connections get split between the two.
To make `bind_to_port()` fail on a port in use, replace the socket options with `set_socket_options()`.
`set_socket_options()` replaces the defaults entirely. Setting `SO_REUSEADDR` on Linux and macOS keeps the "restarted server can bind again right away" behavior.
> **Note:**`SO_REUSEADDR` alone isn't enough on Windows. Two sockets that both set it can bind to the same port, so use `SO_EXCLUSIVEADDRUSE` instead.
> **Note:** To auto-pick a free port, see [S17. Bind to any available port](../s17-bind-any-port). Under the hood, that's just `bind_to_any_port()` + `listen_after_bind()`.
The server rejects HTTP methods it does not know with `400 Bad Request`. To accept an extension method, such as the WebDAV methods of RFC 4918 (`PROPFIND`, `PROPPATCH`, `MKCOL` and friends) or UPnP's `SUBSCRIBE`, register a handler with `CustomRoute()`. Registering the handler is what makes the server accept the method.
Patterns work the same way as they do for `Get()`. Regular expressions and path parameters are both available.
## Advertise your methods with OPTIONS
A WebDAV client asks the server about its capabilities with `OPTIONS` before doing anything else. cpp-httplib generates neither the `DAV:` header nor `Allow`, so return them yourself. Forget this and clients will turn you away even though your `PROPFIND` works.
- The method name has to be a valid HTTP method token (RFC 9110), and it must be registered before you call `listen()`
- `GET`, `HEAD`, `POST`, `PUT`, `DELETE`, `CONNECT`, `OPTIONS`, `TRACE`, `PATCH` and `PRI` cannot be registered here. Use the dedicated methods for those
- A rejected registration makes `is_valid()` return `false` and `listen()` fail, so the server never starts holding a handler that would never run
- Static file serving and WebSocket upgrades stay `GET`/`HEAD` only
> **Note:** cpp-httplib takes you as far as routing the method. If you want to call it WebDAV, generating the `207 Multi-Status` XML, interpreting the `Depth` header and managing locks are all yours to implement. The protocol itself lives outside the library.
> For the basics of registering handlers, see [S01. Register GET / POST / PUT / DELETE handlers](../s01-handlers).
title: "T02. Control SSL Certificate Verification"
order: 43
order: 44
status: "draft"
---
@@ -51,3 +51,5 @@ On most Linux distributions, root certificates live in a single file like `/etc/
> The same APIs work on the mbedTLS and wolfSSL backends. For choosing between backends, see [T01. Choosing between OpenSSL, mbedTLS, and wolfSSL](../t01-tls-backends).
> For details on diagnosing failures, see [C18. Handle SSL errors](../c18-ssl-errors).
> For TLS configuration on a WebSocket client (`wss://`), see [W05. Configure TLS for wss:// Connections](../w05-websocket-tls).
> For mTLS with a WebSocket client (`wss://`), see [W05. Configure TLS for wss:// Connections](../w05-websocket-tls).
## Read client info from a handler
To see which client connected from inside a handler, use `req.peer_cert()`. Details in [T05. Access the peer certificate on the server](../t05-peer-cert).
title: "W01. Implement a WebSocket Echo Server and Client"
order: 51
order: 52
status: "draft"
---
@@ -36,6 +36,7 @@ The `read()` return value is a `ReadResult` enum:
- `ReadResult::Text`: received a text message
- `ReadResult::Binary`: received a binary message
- `ReadResult::Fail`: error, or connection closed
- `ReadResult::Timeout`: a read timeout you set with `set_read_timeout()` elapsed with nothing received; the connection is still open. The compile-time default timeout closes the connection and is reported as `Fail` instead — see [W06. Set Timeouts](../w06-websocket-timeouts)
## Client: talk to the echo server
@@ -85,4 +86,4 @@ svr.new_task_queue = [] {
See [S21. Configure the thread pool](../s21-thread-pool).
> **Note:** To run WebSocket over HTTPS, use `httplib::SSLServer` instead of `httplib::Server` — the same `WebSocket()` handler just works. On the client side, use a `wss://` URL.
> **Note:** To run WebSocket over HTTPS, use `httplib::SSLServer` instead of `httplib::Server` — the same `WebSocket()` handler just works. On the client side, use a `wss://` URL. For CA and client certificate configuration, see [W05. Configure TLS for wss:// Connections](../w05-websocket-tls).
@@ -75,6 +75,6 @@ The counter is reset whenever `read()` consumes an incoming Pong frame, so this
`max_missed_pongs` defaults to `0`, which means "never close the connection because of missing pongs." Pings are still sent on the heartbeat interval, but their responses aren't checked. If you want unresponsive-peer detection, set it explicitly to `1` or higher.
Even with `0`, a dead connection won't linger forever: while your code is inside `read()`, `CPPHTTPLIB_WEBSOCKET_READ_TIMEOUT_SECOND` (default **300 seconds = 5 minutes**) acts as a backstop and `read()` fails if no frame arrives in time. Think of`max_missed_pongs`as the knob for detecting an unresponsive peer**faster** than that.
On the server side, even with `0`, a dead connection won't linger forever: while a handler is inside `read()`, `CPPHTTPLIB_WEBSOCKET_SERVER_READ_TIMEOUT_SECOND` (default **300 seconds = 5 minutes**) acts as a backstop. A client has no backstop of its own — it waits forever unless you set a read timeout — so there`max_missed_pongs`is what notices an unresponsive peer at all. On either side, it is also how you notice one**faster** than that 5-minute fallback.
> For handling a closed connection, see [W03. Handle connection close](../w03-websocket-close).
title: "W05. Configure TLS for wss:// Connections"
order: 56
status: "draft"
---
Client-side TLS configuration for `wss://` (WebSocket over TLS) connections uses almost the same API as `SSLClient`. `ws::WebSocketClient` handles both `ws://` and `wss://` through the same class, so there's no separate class to switch to the way `SSLClient` requires.
Use `set_ca_cert_path()` to point at your own CA certificate. The signature matches `SSLClient`: the first argument is the CA certificate file, the second is an optional CA directory.
To disable certificate verification entirely, use `enable_server_certificate_verification(false)`. For details on that behavior, see [T02. Control SSL Certificate Verification](../t02-cert-verification).
## Presenting a client certificate (mTLS)
`ws::WebSocketClient` has a constructor overload that takes a `PemMemory` struct, letting `wss://` connections present a client certificate.
Passing `PemMemory` to a `ws://` (non-TLS) URL is silently ignored. There's no constructor that reads the cert files directly, so unlike `SSLClient` you always load the PEM into memory yourself before passing it in.
For the full mTLS picture, including server-side setup and use cases, see [T04. Configure mTLS](../t04-mtls).
Set the connection and write timeouts before calling `connect()`. The read timeout can be changed at any time — setting it on an open connection takes effect on the next `read()`.
## Use `std::chrono`
Just like `Client`, there's an overload that takes a `std::chrono` duration directly.
```cpp
using namespace std::chrono_literals;
ws.set_connection_timeout(5s);
ws.set_read_timeout(30s);
ws.set_write_timeout(10s);
```
## What the read timeout means
`set_read_timeout()` applies to a single `read()` call. If no message arrives within that time, `read()` returns `ReadResult::Timeout`: **the connection is still open** and nothing was consumed, so you can send on it and read again. That is what separates it from `ReadResult::Fail`, which means the connection is gone.
This is what lets one thread own a connection in both directions:
```cpp
using namespace std::chrono_literals;
ws.set_read_timeout(100ms);
std::string msg;
while (ws.is_open()) {
auto r = ws.read(msg);
if (r == httplib::ws::Timeout) {
flush_outgoing(ws); // nothing arrived — send whatever is queued
continue;
}
if (r == httplib::ws::Fail) { break; }
handle(msg);
}
```
Without a read timeout, `read()` blocks until a message arrives, so the thread holding the connection never gets to its writes.
Two things to know about `Timeout`:
- It leaves `msg` untouched, and it is non-zero. So `while (ws.read(msg))` is not usable once a read timeout is set — the loop would keep running with the *previous* message still in `msg`.
- It is only reported on a message boundary. If the timeout elapses partway through a fragmented message, that message cannot be resumed and `read()` returns `Fail`.
For connections where long idle periods are normal — waiting on notifications, for example — either leave the read timeout unset, or treat `Timeout` as the no-op it is and keep looping.
## On the server side
A handler's `ws::WebSocket` has `set_read_timeout()` too, and the pattern above is how a handler relays between connections instead of parking in `read()`.
The server default is 300s (`CPPHTTPLIB_WEBSOCKET_SERVER_READ_TIMEOUT_SECOND`) rather than "forever": it is a backstop that reclaims a worker from a peer that has gone silent, since a WebSocket handler holds its worker for the life of the connection.
Because it is a backstop rather than something the handler asked for, it does not surface as `Timeout`. When it elapses, `read()` returns `Fail` and closes the connection, so a handler written as `while (ws.read(msg))` ends the way it always has. Only a timeout the handler set itself with `set_read_timeout()` comes back as `Timeout`.
> Unresponsive-peer detection via Ping/Pong is a separate mechanism. See [W02. Set a WebSocket Heartbeat](../w02-websocket-ping) for details.
## How this differs from `Client`
For `Client`'s timeout configuration, see [C12. Set Timeouts](../c12-timeouts). The behavior and API are nearly identical, but `WebSocketClient` has no equivalent to `set_max_timeout()` for capping the whole request — once connected, the connection stays open for as long as you keep calling `read()`.
A check inside the handler runs after the handshake has completed. To refuse the connection with an HTTP status such as 401 before it is upgraded, use `set_pre_request_handler()` instead. It also runs for WebSocket routes. See [S11. Authenticate per route with a pre-request handler](../../cookbook/s11-pre-request).
## Using WSS
WebSocket over HTTPS (WSS) is also supported. On the server side, just register a WebSocket handler on `httplib::SSLServer`.
サーバーは知らないHTTPメソッドを`400 Bad Request`で弾きます。RFC 4918のWebDAVメソッド(`PROPFIND`、`PROPPATCH`、`MKCOL`など)やUPnPの`SUBSCRIBE`のような拡張メソッドを受け付けたいときは、`CustomRoute()`でハンドラを登録してください。登録したことがそのまま「このメソッドを受け付ける」という意味になります。
// The connection survived the heartbeat exchange without a TLS-session race.
EXPECT_TRUE(cli.is_open());
cli.close();
reader.join();
}
#endif // CPPHTTPLIB_SSL_ENABLED
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.