Lookups in objects are linear, as for ordered_json. Objects with 128
members or more now get a hash table after parsing (open addressing; the
first of duplicate keys is kept, as for the linear search), so that
operator[], at(), find(), contains(), count(), value(), and JSON pointers
take constant time on average in them; the idea of switching to a hash
table for large objects is Boost.JSON's. The parser notes such objects when
it closes them (out of line, so that the parse loop only has a call for
it), and the object node keeps the number of its table.
Looking up each key of an object with 10,000 members: 59.8 ms -> 0.16 ms.
Parsing (json_document::parse, best of 7, separate processes): most files
within 1%; canada +5%, mesh.pretty +3%, citm +3%.
Tests: objects with 127, 128, 129, and 10,000 members (escaped, empty,
and duplicate keys, missing keys, comparisons), nested large objects, and
documents reused with read().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Hold the UTF-8 lookup tables in std::array, compute the length of a
sequence without nested conditionals, and use std::array in the tests.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Long runs of string bytes are scanned 16 at a time with NEON (AArch64, with
GCC and Clang) and SSE2 (x86-64): both belong to the baseline instruction
sets. A signed compare with 0x20 finds control characters and non-ASCII
bytes at once. Keys keep 16 table checks before the vector loop (their
lengths repeat from record to record, so the branches predict well);
string values have 8, as their lengths vary more.
Non-ASCII text is validated 16 bytes at a time with the "lookup4" check of
simdjson (J. Keiser and D. Lemire, "Validating UTF-8 In Less Than One
Instruction Per Byte", 2021): with NEON, and on x86-64 with SSSE3 if
JSON_VIEW_USE_SSSE3 is defined (SSSE3 is not part of x86-64, and the code
must not depend on the flags of a translation unit). JSON_VIEW_NO_SIMD
selects the portable code. The vector code sits in
detail/view/simd.hpp; the same input is accepted either way.
json_document::parse, best of 7 runs in separate processes (M1 Max):
poet.json (CJK text) -72%, random.json -25%, twitter.json -22%,
gsoc-2018.json -20%, semanticscholar -19%, github_events -11%,
apache_builds -9.5%, canada/citm -5/-6%; lottie +4%, tree-pretty +2.5%.
Tests: every two-byte sequence and three- and four-byte sequences with
continuation bytes at the edges of their ranges, at every offset around
the vector blocks of keys and values, cut short, and long runs of text
with a damaged byte, against json::accept and json::parse. CMake builds
the parser tests again with JSON_VIEW_NO_SIMD, and on x86-64 with
JSON_VIEW_USE_SSSE3 and -mssse3; the macros are documented.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
A workflow runs compare.py with pinned downloads on GitHub-hosted
Ubuntu runners and shows the results as the job summary and as an
artifact: started by hand (workflow_dispatch: x86-64 or AArch64, GCC or
Clang), or when a pull request gets the label "benchmark" (both
architectures, GCC). The label trigger gives numbers before the
workflow is on the default branch, which workflow_dispatch needs.
Shared runners are noisy, so the numbers show where json_view stands on
another architecture; published numbers still need a quiet machine.
compare.py takes the CPU name from lscpu where /proc/cpuinfo has none
(AArch64 Linux), and falls back to the architecture. Checked in Linux
containers (AArch64, Clang 15 and GCC 9, offline with the pinned
archives).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
compare.py --download now checks the SHA-256 of the yyjson 0.13.0,
simdjson 4.6.11, and Boost 1.92.0 archives. It unpacks each archive
once (Boost's directory is boost_1_92_0) and, where Python supports it,
with the 'data' filter.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
tests/benchmarks/json_view/ holds the comparison with other libraries,
which is not built by CMake or run by CI:
- bench_view.cpp: parse, traverse, select, and dump of twitter,
citm_catalog, canada, jeopardy, a single tweet, and a JSON-RPC request,
with json_view, yyjson, simdjson (DOM and On-Demand), Boost.JSON, and
json::parse; all engines must agree on every document before anything
is timed, and run interleaved in every round
- bench_corpus.cpp: parse, traverse, and dump of any list of files
- compare.py: builds both against include/ with the libraries of the
system (or pinned downloads), runs them, and writes the results with
what is needed to reproduce them (date, commit, CPU, OS, compiler,
flags, library versions) to results/<date>-<host>.md and .csv; only the
Python 3 standard library is used
- README.md: how to run it, what is measured, and which features the
engines have, so the numbers can be read correctly
Boost.JSON is optional (JSON_VIEW_BENCH_BOOST).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The parser records where the integer digits, the fraction digits, and the
exponent of a float token are. For doubles with at most 19 digits, the
value is now read from that layout: the digits eight at a time, without
scanning the token, and rounded with Clinger's fast path where both
operands are exact, else with the Eisel-Lemire algorithm (which needs no
fallback for up to 19 digits). Both round correctly, so the values are
those of parse(); other tokens and types keep the library's conversion.
get<double>(), materialize(), dump(), and comparisons use it. Traversing
canada.json (111,000 floats, every number converted): 0.53 -> 0.86 GB/s.
Tests add tokens around the limits (19 and 20 digits, 2^53, 10^22) to the
bit-for-bit comparison with parse().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Separate the comparison of discarded values from the other types, so
that the conditional chain has no repeated branch bodies, and mark
the deliberate comparisons of views with empty containers in the
tests.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- API pages for operator== and operator!= of basic_json_view, linked
both ways with the basic_json pages
- the feature page and the class overview list the comparisons
- the examples show when the view helps: detecting a changed document
without building json values
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_view gains operator== and operator!= with other views and with
basic_json values. Two views are equal if the values parse() would
produce for them are equal by basic_json's operator==: numbers compare by
value across their types, and objects by their members, with duplicate
keys resolved as parse() resolves them (the last value, at the position of
the first key). Objects are compared in member order if the object type
keeps an order (ordered_json), by key otherwise, as basic_json does.
Discarded views compare as discarded basic_json values do, which follows
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON. Nothing is materialized except
single numbers, and the walk is iterative.
Tests compare the results for pairs of 1,200 generated documents (also
written differently: sorted keys, canonical numbers) with those of
basic_json, for json and ordered_json, plus numbers, duplicate keys,
member order, discarded values, and 100,000 levels of nesting.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The output buffer initializes its members in the initializer list, and the
escaping has no nested conditional operators; the test marks a fixed seed.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- API pages for dump, number_format, and operator<< of basic_json_view,
linked both ways with the basic_json pages
- the feature page describes document order and number_format::source
- the examples show when the view helps: forwarding part of a message
and writing numbers exactly as they were read
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_view::dump(indent, indent_char, ensure_ascii, number_format)
writes the text of a value as ordered_json::parse(text).dump() writes it
for the same arguments: members in document order (all of them, should a
key occur more than once), strings escaped by the same rules and with the
library's scanning kernels, floats with the library's conversion, and
integers copied from the source, where they are canonical except "-0".
With number_format::source, numbers are copied as they appear in the
source ("1.50", "1E2", "-0", all digits of long integers). operator<<
takes the indentation from the stream width, as for basic_json.
The writer (detail/view/serializer.hpp) writes through a raw pointer into
a string sized from the source extent of the value, and walks the index
iteratively, so the nesting depth is limited by memory only.
Tests compare the output of 2,000 generated documents with
ordered_json::dump() for several indentations and ensure_ascii, strings
with every kind of escape, numbers (5,000 random doubles, float as
number_float_t), duplicate keys, 100,000 levels of nesting, and streams.
ViewDump joins the benchmarks.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
get<T>() of arithmetic types is inlined down to the conversion, so that
its checks of the node kind merge with those of the caller, and reading
an integer needs no call. Traversing every value: citm_catalog -6%,
marine_ik -5%, numbers and twitter -3%, mesh -2.5%, canada -1% (and
more above the float conversion from the digit layout: citm_catalog
-14%, marine_ik -11%).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
get_string() and number_token() return braced lists; the test compares
floats by their bit patterns instead of with memcmp, uses std::any_of, and
marks a fixed seed, a default member initializer (needed by GCC's
-Weffc++), and a string search.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- API pages for get, get_to, get_string, number_token, and value of
basic_json_view; JSON pointer overloads of operator[], at, and
contains; links both ways with the basic_json pages
- the feature page describes which conversions copy nothing
- the examples show when the view helps: strings without copies, numbers
exactly as written, and paths into a large text
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_view gains get<T>(), get_to(), value() with keys and JSON
pointers, and operator[], at(), and contains() with JSON pointers, plus
two functions basic_json has no counterpart for:
- get_string(): the string without a copy (a string_view into the source,
or into the decoded strings for strings with escapes)
- number_token(): the text of a number as it appears in the source
get<T>() converts arithmetic types, strings (also string_view_t),
std::nullptr_t, std::vector, maps with string keys, and views directly;
floats are converted from the digit layout recorded by the parser with the
library's conversion chain, so the values are bit-identical to parse().
Other types, including user types with from_json(), go through
materialize().
The exceptions are those of basic_json, message included. Where const
basic_json has undefined behavior (a missing key or an index out of range
with operator[] and a JSON pointer), the result is a discarded view;
value() returns the default wherever basic_json catches out_of_range, and
contains() never throws. Array indices of JSON pointers follow
json_pointer's rules (parse_error.106/109, out_of_range.404/410).
Tests compare the conversions of 2,000 generated documents, 20,000 float
tokens (double and float, bit for bit), and every JSON pointer of 1,000
documents with basic_json, and the exceptions for malformed pointers.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
detail::json_pointer_access returns the reference tokens of a pointer, so
that code resolving pointers without a basic_json value (such as the
zero-copy view) does not have to parse to_string() again. No change in
behavior.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Marks the default initializer of the item's index string (needed by GCC's
-Weffc++) and, in the test, an escaped literal and a comparison of find()
with end(), which is what the test is about.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- API pages for operator[], at, front, back, find, contains, count,
begin, end, cbegin, cend, items, and type_name of basic_json_view,
linked both ways with the basic_json pages
- the feature page and size() describe document order and duplicate
keys
- the examples show when the view helps: reading a few fields of a large
text, probing optional members, and members in source order
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_view gains the read-only access functions of basic_json:
operator[] and at() with keys and indices, front(), back(), find(),
contains(), count(), begin()/end(), items() (with structured bindings from
C++17 on), and type_name(). They throw the exceptions (ids and messages)
that the const functions of basic_json throw; where basic_json has
undefined behavior (operator[] with a missing key or an index out of
range, front()/back() of an empty container), the view returns a
discarded view or throws invalid_iterator.214.
Objects are iterated in document order, and all members are visited. With
duplicate keys, lookups find the first member, so that a lookup can stop
at the first match; parse() keeps the last value. Keys of up to 16 bytes
are compared with two overlapping loads instead of memcmp, and most keys
are rejected by their length alone, from the index.
The iterators and items live in detail/view/iterator.hpp, the lookups in
detail/view/lookup.hpp. Tests compare every element and member of 2,000
generated documents with ordered_json, keys of every length around the
load sizes, the exceptions against const basic_json, and the iterators.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json::type_name() now calls detail::value_type_name(value_t), so
that code which reports types without a basic_json value at hand, such as
the zero-copy view, uses the same names in its exception messages. No
change in behavior.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
A new section describes the 16-byte node: its fields, how integers,
floats, and object members are stored, how views navigate without
pointers, and a worked example. The feature page and the pages of
basic_json_document and node_count link to it where they mention the
16 bytes.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The 4 GiB limit and the fallback for an input that parse() accepts but
the view rejects (a bug) are excluded from the coverage; shrink_to_fit()
of an empty document is tested.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- the input dispatch takes byte ranges by const reference and reads the
size once (which also settles a finding of the static analyzer); input
adapters are taken by value
- the classification of inputs keeps its nested conditional operators, a
constant expression of C++11 (NOLINT)
- the test's C arrays, fixed seed, and escaped literals are marked, as in
the other tests
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
benchmarks_view.cpp adds ViewParse, ViewRead (a reused document),
ViewParseIndented, ViewAccept, and ViewMaterialize on the files of
ParseString, so that each row can be read against the json::parse row of
the same file; benchmarks.cpp gains Accept (json::accept) as the
counterpart of ViewAccept. The view benchmarks are built only if the
header directory has json_view.hpp, so that older versions can still be
benchmarked.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- API pages for basic_json_document and basic_json_view, one per member,
and for the four aliases, each with an example
- features/json_view.md: the problem the view solves, ownership and
lifetime, what matches basic_json::parse() and what differs, and when
to choose json, ordered_json, SAX, or the view
- the examples show why one would use the view, not only how: borrowed
vs. owned input, reading a few fields and materializing one subtree,
reusing a document across many messages
- registered in the mkdocs navigation, llms.txt, the docset, the
exceptions page (out_of_range.416), architecture.md, the integration
page, and the README; the yyjson credit is added to the README and
license.md
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
fuzzer-parse_json_view.cpp checks for every input that
json_document::accept agrees with json::accept, that an accepted input
materializes to the value json::parse returns, and that a rejected input
makes both throw the same exception with the same message. It is built
like the other fuzzers (tests/Makefile, and the root Makefile's
fuzz_testing_json_view target, which starts from the JSON test corpus) and
listed in tests/fuzzing.md.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The nlohmann.json module includes json_view.hpp in its global module
fragment and exports basic_json_document, basic_json_view, and the four
aliases next to basic_json, json, and ordered_json. features/modules.md
lists them, and tests/module_cpp20 parses a document, so that a missing
export fails the ci_module_cpp20 job.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The new single header goes through the same checks and install steps as
json.hpp and json_fwd.hpp:
- cmake/ci.cmake: ci_test_amalgamation regenerates, formats, and compares
json_view.hpp as well
- check_amalgamation.yml: the pull request check does the same; it runs
develop's tools, so it needs config_json_view.json on develop first
- meson.build: installs single_include/nlohmann/json_view.hpp
- gen_bazel_build_file.cmake, BUILD.bazel: json_view.hpp joins the
single-header target; the glob of the other target already covers the
new headers
- labeler.yml: an "aspect: json_view" label for the header, its
detail/view headers, tests, and documentation
The CMake install rules for include/ and single_include/, the REUSE
catch-all, and Package.swift cover the new files without changes.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
unit-json_view.cpp: type queries, size, and empty against basic_json;
materialize() against parse() (json and ordered_json, generated documents,
duplicate keys, 100,000 levels of nesting, parent pointers with
JSON_DIAGNOSTICS); parse errors and their messages equal to parse() for
malformed inputs and all option combinations; NUL and BOM; borrowed and
owned inputs (strings, C strings, literals, vectors, string_view, streams,
wide strings, parse_copy, and iterator ranges over pointers, vectors,
strings, and lists); reuse with read(); moves; shrink_to_fit() of the index
and of the decoded strings; source offsets.
unit-json_view_macros.cpp includes the header without
JSON_TEST_KEEP_MACROS, as users do: the view must not depend on the macros
json.hpp undefines, and must not leak its own.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The public classes of the zero-copy view (#5295), in the new header
<nlohmann/json_view.hpp>:
- basic_json_document<BasicJsonType>: parse (borrowing contiguous byte
inputs, owning rvalue strings, streams, and other inputs), parse_copy,
accept, read, root, is_discarded, source, owns_source, node_count,
memory_usage, shrink_to_fit
- basic_json_view<BasicJsonType>: type and the is_* queries, size, empty,
materialize (the value parse() would produce, built by the same SAX
handler), source_offset
- the aliases json_document, json_view, ordered_json_document, and
ordered_json_view
A parse error throws the exception basic_json::parse would throw for the
same input: the library parser is run on the failing input, so messages,
positions, and exception ids are the same. Inputs of 4 GiB or more are
rejected with out_of_range.416.
The single header single_include/nlohmann/json_view.hpp keeps including
json.hpp; make amalgamate, check-amalgamation, include.zip, and release
handle it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Six members of string_ref (length, begin, end, operator[], operator!=,
and operator<<) were not reached before C++17, where string_ref is
std::string_view.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Since whole blocks of eight digits are read directly, parse_upto8() only
gets fewer than eight digits; its eight-digit case was dead code.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
A block of eight digits of a number token lies inside the input (the
digits were counted while scanning, or are recorded in the digit
layout), so parse_upto19() reads it without the bounds check of the
last, partial block. Traversing canada.json: -11% instructions, -6%
cycles.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
json::parse ends the input at a NUL only between values (where it does
at all); inside a string, a NUL is a control character that must be
escaped. The view reported it as a missing closing quote. (Only the
error code differed: the exception comes from the library parser.)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Numbers of more than 19 digits at the boundary of the largest double
are decided by the locale-aware fallback of the overflow check.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- tables as std::array; the frames of the first 64 levels stay a C array
(not initialized on purpose, NOLINT)
- \u escapes are decoded with the library's hex_codepoint() instead of a
second table
- the parse failure is private, with an accessor; the special member
functions of the builder are all declared
- no nested conditional operators; explicit parentheses; a repeated
branch body merged; auto for casts
- the test's C arrays, fixed seed, and escaped literals are marked, as in
the other tests
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
With JSON_NOEXCEPTION, NLOHMANN_VIEW_THROW(e) was std::abort() alone, so
the parameters of the functions that build the exceptions were unused, a
warning that the builds with -Werror turn into an error. The exception is
now evaluated before std::abort(); the program ends anyway.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The builder parses JSON text in one pass into the node index: strings and
numbers stay in the source (escaped strings are decoded into an arena),
integers are converted while their digits are in cache, and floats keep
their digit layout for a later conversion. It accepts exactly what
json::parse accepts, for every combination of comments and trailing commas,
with and without a terminating NUL, and with JSON_STRICT_NUL_HANDLING.
Parse state lives in a local cursor whose address never escapes, so that it
stays in registers; out-of-line helpers (errors, regrowth, escapes,
comments) are members of the builder and get the positions they need. The
value dispatch is expanded once for array elements and once for member
values. Literals are compared with memcmp and words read in a fixed byte
order, so nothing depends on the platform's byte order. Error messages come
with the public classes.
Tests (unit-json_view_builder.cpp): accept/reject and values against
json::parse for handwritten, generated, and damaged documents under all
option combinations, from std::string and from exact-size buffers (no read
past the input under AddressSanitizer), deep nesting up to 100,000 levels,
NUL/BOM/whitespace cases, and the test-suite files.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Internal parts of the zero-copy view (#5295), under detail/view and not
included by json.hpp, so that users of json.hpp compile nothing of it:
- macro_scope.hpp/macro_unscope.hpp: the few macros the view needs, under
its own prefix (json.hpp undefines its own at its end); the throw macro
honors JSON_NOEXCEPTION and JSON_THROW_USER like JSON_THROW
- string_ref.hpp: std::string_view from C++17 on, else a small stand-in
- node.hpp: the 16-byte node of the index; its kinds are value_t values
(checked by a static_assert)
- document_data.hpp: the storage of a parsed document (node array, decode
arena, owned input)
- scan.hpp: string and digit scanning with unrolled checks at fixed offsets
(after yyjson) and the library's SWAR and UTF-8 checks, independent of the
byte order
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
json.hpp undefines JSON_STRICT_NUL_HANDLING and
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON at its end. Code that builds on
the library after it, such as the planned json_view.hpp, reads them from
detail::abi_config instead. The constants live in the ABI namespace, which
already encodes both settings, so they always match the basic_json in use.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
get_codepoint() read the four hex digits of a \u escape with four calls
to get(), each classified by a chain of range comparisons. For contiguous
input, get_codepoint_bulk() now decodes them with one lookup per byte
(hex_codepoint() in string_scan.hpp, after yyjson's read_hex_u16): a
256-entry table maps a byte to its value, or 0xFF for anything else, and
an invalid digit shows in the OR of the four values. It then skips the
four bytes and updates the position counters as four get() calls would.
If a digit is invalid or fewer than four bytes are left, it changes
nothing and the existing loop runs, so errors are reported with the same
message and position as before.
json::parse, best of 5 runs in separate processes (M1 Max): the escaped
twitter.json (every non-ASCII character as \u) -13.6%, all other files
within 0.3%.
Tests compare the contiguous and the streaming path (value or exception
message) for valid escapes, surrogate pairs, truncated and invalid digits
at every position, and 3,000 seeded random escapes.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
find_string_special() and find_ascii_copyable_run() test eight bytes at a
time, but located the stopping byte inside a word with a byte loop. The
lowest flagged byte of the SWAR tests is always a true hit (the borrows of
the subtractions can only flag bytes above one), so its index is now the
trailing-zero count of the mask; words are read in little-endian order on
every platform, so this does not depend on the byte order.
scalar_string_bulk_run() validates a run of multi-byte UTF-8 sequences one
after another instead of searching for the next special byte in between,
which helps text in non-Latin scripts.
The kernels serve the lexer's contiguous fast path, the serializer, and the
binary formats. New tests compare all three with byte-by-byte reference
scans on 100,000 generated buffers at three alignments; the portable
fallback of count_trailing_zeros() was checked against the builtin.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Float tokens that Clinger's fast path cannot convert (e.g. the 17-digit
coordinates of canada.json) went to strtod unless std::from_chars was
available. It is not used in C++11/14, and not with libc++, which does not
define __cpp_lib_to_chars. The Eisel-Lemire algorithm (after fast_float's
compute_float) now converts them with integer arithmetic, correctly rounded
for any token with at most 19 significant digits. Longer tokens are
truncated; the result is used if w and w + 1 round alike, else strtod
decides as before. Overflow still yields infinity (out_of_range.406).
The table of powers of five (fast_float's) lives in pow5_table.hpp; a unit
test recomputes every entry with big-integer arithmetic. Further tests:
known values generated with Python (whose float() is correctly rounded),
200,000 round trips through to_chars, and the 128-bit multiplication and
leading-zero count against big-integer references (both with and without a
128-bit type). Checked against strtod on 6.5 million tokens, among them
60,000 exact halfway cases: no difference.
json::parse on canada.json: -8.6% (C++11), -7.6% (C++17, Apple clang);
other files unchanged. Compile time of a TU including json.hpp: +0.7%.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
lexer::convert_number() converted float tokens with std::from_chars (when
available), Clinger's fast path, and the locale-aware strtod fallback, all
as lexer members. They are now free functions in number_parse.hpp:
- convert_float_fast(): std::from_chars, then Clinger's fast path, skipped
when the mantissa has too many significant digits
- convert_float_locale_aware(): strtof/strtod/strtold with the decimal point
of the current locale, retried when the locale changed (#5198)
so that other code converting JSON number tokens gets the same values. No
change in behavior; the lexer no longer includes <clocale> and <cstdlib>.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
A header that builds on json.hpp (such as the planned json_view.hpp) must
not inline json.hpp: its single-header version would contain a second copy
of the library, and that copy would change with every library change.
The optional config key "external" lists include paths that are kept as
#include directives. Only the first directive per path is kept; repeated
ones are commented out, as the tool already does for inlined headers.
config_json_view.json uses it for json_view.hpp; the existing configs do
not set it, and json.hpp and json_fwd.hpp regenerate byte-identically.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
#5597 was merged before all of its CI jobs had run, and two of them fail
on develop now, and so on every pull request:
- ci_clang_tidy: cert-err33-c for the two std::setlocale(LC_NUMERIC, "C")
calls whose result was discarded. Check the result, like the other
resets in the file.
- ci_test_standards_gcc (20) with GCC 16: -Wnoexcept for the two parser
callbacks, which cannot throw but were not declared noexcept.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Look up the locale decimal point at conversion time, not lexer construction
The lexer read localeconv()->decimal_point once in its constructor and wrote
that character into token_buffer in place of '.'. The strtod fallback then
used the locale current at conversion time, so an LC_NUMERIC change in
between (parser callback, SAX handler, another thread) truncated the value
in release builds and fired the endptr assertion in debug builds.
token_buffer now always holds '.'. Only the strtof/strtod/strtold fallback
depends on the locale: it looks up the decimal point right before the call,
restores '.' afterwards, and repeats the conversion if the locale changed in
between. As a side effect, std::from_chars and Clinger's fast path now also
apply under locales whose decimal point is not '.'.
Fixes#5198
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Stop the strtod retry loop when the decimal point is unchanged
convert_float_locale_aware() repeated the conversion until strtod
consumed the whole token, assuming an early stop can only mean a locale
change. Under a locale whose decimal point is not a single character
(e.g. the two-byte U+066B of ar_EG.UTF-8, ar_SA.UTF-8, or fa_IR.UTF-8,
all available on macOS), the in-place substitution can never succeed,
so parsing any float that reaches the strtod fallback (for example
3.14159265358979323846 at C++11) hung forever. Before this branch, the
same input was truncated.
Retry only if the decimal point changed since the previous attempt;
otherwise keep the value strtod parsed so far, as before. Add a test
that parses such numbers under a multi-byte decimal point locale; it
hangs without this change.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix -Weffc++ errors in the #5198 locale test
GCC's -Weffc++ (an error in ci_test_gcc and ci_test_standards_gcc)
rejected LocaleSwitchingSax: it has a pointer data member but does not
declare its copy operations, and its vectors are not initialized in the
member initializer list. Store the locale name as a std::string and give
the vectors brace initializers, like SaxEventLogger in
unit-deserialization.cpp.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Name the key type when rejecting non-string CBOR/MessagePack map keys
CBOR and MessagePack allow map keys of any type, but JSON object keys
are always strings, so such maps are rejected. The error so far was the
one for a malformed string (e.g. "expected length specification
(0xA0-0xBF, 0xD9-0xDB); last byte: 0xC0" for a nil key), which does not
tell the user what went wrong. Report the type of the key instead:
syntax error while parsing MessagePack object key: only string keys
are supported, but found nil; last byte: 0xC0
The exception id (parse_error.113) and type are unchanged. Malformed
string keys and a missing key keep their previous messages. Document
the restriction on the CBOR and MessagePack pages.
Refs #2766, #3381
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Point the MessagePack key note to the spec's profile section
The note linked to "Serialization: type to format conversion", which says nothing about key types. Restricting map keys to strings is only mentioned in the "Profile" section (under "Future discussion") as an example of a JSON-compatible profile, so link there and describe it as such instead of as a permission.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The "Size above uint32" tests for arrays and objects fake a container
size of 2^32 and expect to_msgpack() to throw out_of_range.412. But
to_msgpack(j) first reserves binary_reserve_hint(j) bytes, which is
size + 1 for arrays and 2 * size + 1 for objects, i.e. 4 or 8 GiB.
Linux and macOS overcommit, so the reservation succeeds; on Windows it
throws std::bad_alloc before the size check is reached (seen with
msvc-vs2026 Debug x64 on the object test).
Write into a caller-owned vector instead, so nothing is reserved.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
#5559 added a 224-character line to cbor_tag_handler_t.md; the
documentation style check allows at most 160.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The line added in #5559 exceeded the 160-character limit enforced by
docs/mkdocs/scripts/check_structure.py, breaking the documentation build.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The tests added by #5515 fail two ways on develop:
- clang-tidy reports the size() overrides of huge_string and huge_binary
(readability-convert-member-functions-to-static) and the non-const
test value (misc-const-correctness); mark them like the #5584 types
- clang with libstdc++ 10 cannot compile the file for C++17: the
std::filesystem::path conversion considered for huge_string, a class
derived from std::string, is ambiguous. Guard it with
JSON_TEST_BEYOND_UINT32_STRING, which #5584 introduced for the same
reason, and define that macro before both test blocks.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add BON8 support
Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format
that uses the byte values that cannot begin a UTF-8 character as type
markers, so strings need no length prefix. It is the most compact of the
supported binary formats on the benchmark files.
The reader is non-recursive like the other binary readers. A string ends
at the first byte that cannot continue it, so the reader hands the one or
two bytes it reads past a string back to the value that follows. The
writer produces the canonical representation of the specification, except
for NFC normalization; its output is identical to that of the reference
implementation (HikoGUI) on all files of the test data.
The round-trip tests need the .bon8 files of json_test_data 3.2.0.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Address review comments
- Reuse detail::validate_one_utf8 to check strings in to_bon8; the error
now names the first byte of the invalid sequence.
- Document that to_bon8 leaves bytes in the output adapter on an
exception, and that string_open is only an output of write_bon8_marker.
- Explain why the pushback buffer of the BON8 reader cannot overflow.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Select the BON8 float prefix by type
get_bon8_float_prefix only depends on the type of its argument, so make
the type a template parameter instead of passing an unused value.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Rename a test variable that Flawfinder mistakes for read()
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix the BON8 CI failures
- compare the float in write_bon8_float with number_float_t constants,
so GCC does not warn about a float-to-double conversion
- mark check_bon8_utf8's context as used when exceptions are disabled
- choose the compact float prefix in a helper rather than with nested
conditional operators (clang-tidy)
- use auto for the cast in the BON8 integer reader (clang-tidy)
- write the int32 minimum test values as long long literals (MSVC C4146)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Amalgamate
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Read BON8 strings in bulk from contiguous input
- copy the valid UTF-8 of a string in one step when the input is
contiguous (twitter.json is read in 1.68 instead of 2.52 ms,
jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack)
- share the new valid_utf8_prefix() with the writer's UTF-8 check, which
now skips ASCII 8 bytes at a time
- let the fuzzer check that contiguous and stream input give the same
value or error, and test both paths in the unit tests
- clarify that a second 0xFF after a string is an empty string
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Link the BON8 functions from the other binary format pages
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Name the bulk scan flag after the input, not BON8
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Read BSON keys in bulk from contiguous input
BSON keys (and array indices) are C-style strings, which were read byte
by byte. For contiguous input they are now read up to their \x00-byte in
one step, using the same bulk_scan flag as BON8 strings: twitter.json is
read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of
3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys
are almost all one-digit array indices, takes 2 % longer.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix the BON8 CI failures of the bulk-read tests
- skip the contiguous-versus-stream tests of BON8 strings and BSON keys
when exceptions are disabled: they catch the parse errors of invalid
input, and without exceptions the library aborts instead
- use static_cast for the int64 test value (google-readability-casting)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move the explicit basic_json instantiation into its own test file
Linking test-regression3_cpp20 with clang and MinGW failed with
"relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'",
as test-regression2 did before #5511. The explicit instantiation of
basic_json<> for #4825 compiles every member function, including the
BON8 reader and writer, into that object, and it was already close to
the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1,
C++20).
Give the instantiation a file of its own: unit-regression3 is now
1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file
mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built
for the C++17 standard the regression was about.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Convert the bytes of the BON8 test strings explicitly
The str() helper constructed a std::string from a byte range, which
converts each unsigned char implicitly; -fsanitize=integer reports that
for bytes of 0x80 and above (ci_test_clang_sanitizer).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The tests fake a container size of UINT32_MAX + 1, which does not fit
into a 32-bit std::size_t: MSVC rejects the truncation (C4305/C4309
with /WX), and clang-cl wraps the size to 0 so nothing throws. Guard
them with SIZE_MAX > UINT32_MAX like the tests from #5584.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Throw instead of writing MessagePack lengths beyond UINT32_MAX
MessagePack stores the length of a string, binary value, array, or
object in at most 32 bits. For a larger value, to_msgpack wrote no length
at all, so the output could not be read back. It now throws
out_of_range.412, which BSON already uses for its 32-bit length fields.
The check lives in one function, so each length is written by an
if/else chain that ends in a plain else, without a condition that can
never be false. It is tested with string and binary types that report a
size beyond UINT32_MAX without allocating it, like the BSON tests do.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix the CI failures of the MessagePack length check
- mark to_msgpack_length's value as used when exceptions are disabled
(-Wunused-parameter, misc-unused-parameters)
- put "Exception safety" before "Exceptions" in to_msgpack.md, as the
documentation style check requires
- create the test's string value from its type: constructing it from a
beyond_uint32_string_t considers the std::filesystem::path conversion,
which libstdc++ 10 reports as ambiguous for a class derived from
std::string (clang 13)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Skip the MessagePack string length test for clang with libstdc++ 10
C++17 builds consider the std::filesystem::path conversion for the
string type, and with clang and libstdc++ 10 that conversion is
ambiguous for a class derived from std::string. Creating the value from
its type did not avoid it, since any basic_json with that string type
instantiates the check. The binary and ext cases are still tested there.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Keep the MessagePack string test type and its alias in one block
astyle indented the alias oddly when it had an #ifdef of its own after
the binary alias; declare it right after the string type, in the same
block.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Compare unordered objects by key below the nesting bound
Values nested deeper than the nesting bound are compared without the
call stack, walking both objects entry by entry. Two equal objects of a
type that enumerates its entries in no fixed order - std::unordered_map,
say - can be walked in different orders, so they compared unequal, and
a deep copy compared unequal to its original. std::unordered_map's own
operator== does not depend on the order, which is what applies above the
bound.
Where the keys differ, equality now finds the entry by its key instead.
An ordering, and ordered_map, whose operator== compares its entries in
sequence, still decide by the key.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Test unordered object equality without std::unordered_map
basic_json<std::unordered_map> instantiates std::pair<const string,
basic_json> while basic_json is still incomplete. The standard does not
require std::unordered_map to support that, and libstdc++ 6 to 9 as well
as the EDG front ends of icpc and nvc++ reject it, which broke the build
of unit-comparison on those CI jobs.
The test now uses an object type derived from std::map (which, as the
default object type, works everywhere) whose comparator orders keys
ascending or descending as chosen at construction, and whose operator==
does not depend on the order of the entries - the property of
std::unordered_map the test is about.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Compare the test object type's entries with std::all_of
clang-tidy (readability-use-anyofallof) asked for std::all_of instead of
the loop in unordered_object_t's operator==. The entry type is spelled
out, as C++11 needs typename for base_type::value_type and C++20
reports it as redundant.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add tests for uncovered code paths
Cover code the test suite did not reach, found from the Coveralls report
of develop and a local coverage run of HEAD:
- dump() of every kind of value below the bound of the recursive descent
(pretty-printed objects, binary values, discarded values, scalars), and
flushes of the escape and write buffers mid-string and mid-binary
- the iterative comparison: objects with different keys, containers that
are a prefix of each other, and elements that cannot be ordered, each
both at the top level and below the nesting bound
- SAX handlers that stop at any event, including the end of a nested
container, in the BSON, CBOR, MessagePack, UBJSON and BJData readers
- from_bson/cbor/msgpack/ubjson/bjdata returning a discarded value
through the iterator and pointer overloads
- JSON Patch, diff, merge_patch and update(..., true) on ordered_json
- smaller gaps: get_allocator(), to_ubjson/to_bjdata into a string,
value() with an unresolvable JSON pointer, integer/float comparison
below the integer range and with negative fractions, conversion to a
custom binary type, std::formatter::parse on a spec without '}',
unescape() of a lone '~', and the callback parser's start_array()
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Cover more paths that were thought unreachable
- parse_float_fast() declining malformed or inexact input, called
directly since the lexer only passes well-formed numbers to it
- a UTF-16 high surrogate followed by a unit above the low surrogates
- self-assignment of a const_iterator
- a truncated CBOR string read through non-contiguous iterators
- serializing a long double under the de_DE locale, which undoes the
locale's decimal point and thousands separator
- values read from a binary format carrying no diagnostic positions,
with and without a parser callback
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix the CI failures of the new coverage tests
- declare the self-assignment reference const (misc-const-correctness)
- expect the (/path) prefix that JSON_DIAGNOSTICS adds to the messages
of the failing ordered_json patch operations
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Expect the byte range JSON_DIAGNOSTIC_POSITIONS adds to the patch errors
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Build the expected dump of the nested-object test with +=
clang-tidy (performance-inefficient-string-concatenation) reported the
chain of operator+ calls that assembled the expected indented output.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Compare the BJData and UBJSON test outputs byte by byte
Building a std::string from the byte vector converts each byte
implicitly, which -fsanitize=integer reports for bytes of 0x80 and
above (ci_test_clang_sanitizer).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Coverage reported conditions in the binary writer that can never be
false, and marked the code behind them with LCOV_EXCL. Remove them
instead of excluding them:
- CBOR writes the length of a string, binary value, array, or object
exactly like an unsigned integer, only with another major type. One
function, write_cbor_head(), now writes both, so the integer tests
cover every width and the four excluded 64-bit length branches are
gone.
- A last `else if` whose condition holds for every remaining value
(an unsigned value at most UINT64_MAX, a signed one in the range of
int64_t) is now a plain `else`.
- Whether a signed integer fits into an int64 for UBJSON and BJData is
decided by its type at compile time. Only an integer type wider than
64 bits gets a range check and the high-precision fallback.
- The private get_impl(boolean_t*) was never called.
The UBJSON type prefix 'H' of an optimized container of unsigned
integers beyond the range of int64 was reachable although excluded; it
is tested now.
The output is unchanged.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Write BSON in linear time, without recursing per nesting level
to_bson() had two problems with nested values:
- It recursed once per nesting level, so a value nested deeply enough -
100,000 levels on an 8 MiB stack - exhausted the call stack and
terminated the process, although parse() accepts such values without
complaint.
- BSON prefixes every document and array with its length. The writer
computed that length by walking the entire value below it, again for
every nested document it wrote, which made serializing O(size x depth).
A 200-level document took 30 ms instead of 1.
Both passes are now iterative, and each length is computed exactly once:
- calc_bson_sizes() computes the length of every document and array in
one pass, each from the lengths of its entries, into a table ordered
the way they are written.
- write_bson_document() then writes the document, taking each length from
the table.
Everything observable is unchanged, as a differential test against
develop confirms byte for byte:
- The same bytes are written.
- A key containing U+0000 still throws out_of_range.409 for the same
first key, with the same diagnostics path, before anything is written.
- A document too large for BSON still throws out_of_range.412 before
anything is written.
- A binary subtype above 255 still throws out_of_range.415 after the
same partial output.
Only the enclosing objects and arrays are kept on a stack, so a flat
document allocates nothing for it. Measured against develop (clang -O3,
median of 201 runs): flat objects unchanged, flat arrays 37% faster (the
array length was computed twice), a nested 3,000-object document 2x
faster, a 200-level document 33x faster.
to_bson.md documented the quadratic complexity since #5334; it is linear
again.
Fixes#5392 for BSON, and #5308.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Do not require a default-constructible string_t in the BSON writer
GCC 4.9 and MSVC rejected the test's huge_string_t, which has no default
constructor; develop never default-constructed string_t here either.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Let the BSON index-name helper only fill its output parameter
It returned a reference to the string it filled, so callers held a second
name for index_name. Addresses review feedback.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Make JSON_STRICT_NUL_HANDLING part of the ABI tag
JSON_STRICT_NUL_HANDLING (#5534) changes the bodies of inline functions:
the lexer's handling of '\0' and input_adapter() for char arrays. So
translation units compiled with and without it define the same functions
differently, an ODR violation - the case the ABI tag exists for, as with
JSON_BRACE_INIT_COPY_SEMANTICS (_bics). It now appends _snul to the inline
namespace. The macro is new in 3.13.0, so no existing namespace changes.
Its default moves to abi_macros.hpp, and it is only #undef'd without
JSON_TEST_KEEP_MACROS, as for the other ABI macros. The ABI config tests,
the namespace docs, the macro's docs and the Natvis file cover the new tag.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Amalgamate
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Allocate the deep copy's key scratch space with the provided allocator
The iterative deep copy builds each object's keys in a temporary vector of
key/value pairs before handing them to the object's range constructor. That
vector holds basic_json values, so like the values themselves it now uses
AllocatorType instead of std::allocator.
Also document that AllocatorType covers the JSON values, while most
temporary storage still uses std::allocator.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Count allocate_at_least in the scratch-counting test allocator
From C++23 on, libc++'s containers allocate through allocate_at_least when
the allocator has one. The test allocator inherited it from std::allocator,
so the scratch allocations were not counted and the test failed on Xcode.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* 🐛 fix BSON conformance issue
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* 🐛 fix BSON conformance issue
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* 🐛 reject ill-formed UTF-8 in CBOR/MessagePack/BSON text strings at decode time (#5531)
from_cbor()/from_msgpack()/from_bson() copied the raw bytes of a decoded
text string into the resulting json value without any UTF-8 validation,
even though RFC 8949 §3.1 (CBOR) and the MessagePack/BSON specifications
all require text strings to be valid UTF-8. Malformed input only failed
later, if the value was dump()'d, with a type_error.316 - so the
allow_exceptions=false pattern used specifically to get a discarded
sentinel instead of an exception did not discard this category of
malformed input, unlike every other kind of malformed binary input this
library rejects at decode time (see #5529).
Fix this at the single choke point shared by BSON/CBOR/MessagePack/UBJSON
string reads, binary_reader::get_string(): validate the bytes with the
UTF-8 DFA right after they are read, and report failures the same way as
every other binary_reader error (parse_error.113), so allow_exceptions
and strict discarding behave consistently. get_binary()/binary blob reads
are untouched and still accept arbitrary bytes, since only text strings
are required to be UTF-8.
There were two independent implementations of a UTF-8 validator: the
lexer's streaming scanner, and the serializer's Hoehrmann DFA used by
dump_escaped_impl(). Rather than write a third, the serializer's decode()
function, its utf8d table and the UTF8_ACCEPT/UTF8_REJECT constants are
extracted into detail/string_utils.hpp (a low-level header already
included before both detail/input/ and detail/output/), alongside a new
is_valid_utf8() helper built on the same decode() step. serializer.hpp's
dump_escaped_impl() now calls the shared decode(), so there is exactly
one UTF-8 validator in the codebase; dump()'s exact type_error.316
messages and byte-index reporting are unchanged (see the added
regression-guard test in unit-serialization.cpp).
Claude-Session: https://claude.ai/code/session_01N4RQ1Ahan5YAGbnAQGjZTY
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* ⚡ validate only newly read bytes of binary-format strings
get_string() validated the whole result after each call, but get_bytes()
appends to it and CBOR indefinite-length strings collect all chunks in
the same result, so every chunk re-validated everything read before it.
An input of many small chunks took quadratic time (80000 one-byte chunks,
160 KB of input, took about 7 seconds). Only the newly read bytes are
validated now, which also matches RFC 8949's requirement that every
chunk is valid UTF-8 on its own.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* Support zero-member types in NLOHMANN_DEFINE_TYPE_* macros (#4041)
NLOHMANN_DEFINE_TYPE_INTRUSIVE(Type) and its 11 sibling macros produced
broken code for types with no members to serialize. Invoking a variadic
macro so __VA_ARGS__ is empty is only standard-conforming since C++20,
so a plain __VA_OPT__ fix (as tried in #5142) breaks every pre-C++20
build under -pedantic. Instead, make all 12 macros purely variadic and
dispatch on argument count using a sentinel-padded extension of the
existing NLOHMANN_JSON_GET_MACRO idiom, giving full C++11-C++26 support
with no feature-test gate.
Verified against real GCC 16 and Clang at -std=c++11/14/17/20 with
-pedantic -Werror -Wvariadic-macros: zero regressions in the existing
unit-udt_macro.cpp suite plus 12 new zero-member test cases.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix CI failures in zero-member NLOHMANN_DEFINE_TYPE_* macros
Three issues surfaced on PR #5272's real CI that weren't caught by
local testing against a narrower flag set:
- GCC -Werror=noexcept: the four truly-empty from_json bodies (plain
INTRUSIVE/NON_INTRUSIVE, with and without _WITH_DEFAULT) provably
never throw but weren't declared noexcept; mark them noexcept
explicitly. to_json and the derived-type from_json overloads are
left alone since they genuinely can throw (object assignment /
delegating to the base class's from_json).
- clang-tidy bugprone-macro-parentheses: false positive on the same
8 zero-member bodies (Type/BaseType used purely as declarator
types); suppressed with NOLINTNEXTLINE comments in the same style
already used elsewhere in this file (see NLOHMANN_JSON_SERIALIZE_ENUM).
- MSVC's traditional preprocessor doesn't fully expand
NLOHMANN_JSON_CAT(prefix, NLOHMANN_JSON_TYPE_TAG(...))(...) in one
pass, which broke a pre-existing one-member usage in
unit-regression2.cpp with syntax errors. Wrap all 12 public
dispatcher macros in an extra outer NLOHMANN_JSON_EXPAND(...),
matching the pattern NLOHMANN_JSON_PASTE already uses for the same
MSVC quirk.
Re-verified against real GCC 16 and Clang at -std=c++11/14/17/20 with
-pedantic -Werror -Wvariadic-macros -Wnoexcept, including the exact
files that failed in CI (unit-udt_macro.cpp, unit-regression2.cpp),
against both the modular headers and the re-amalgamated single header.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix clang-tidy misc-const-correctness in unit-udt_macro.cpp
The four zero-member ONLY_SERIALIZE test objects are only ever read
(via to_json), never mutated, so mark them const per clang-tidy.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix derived-type macro dispatch capping members at 62 instead of 63
NLOHMANN_JSON_GET_MACRO resolves 64 positional arguments, with NAME at
position 65. NLOHMANN_JSON_TYPE_TAG dispatches on Type plus the member
list, so it resolves correctly up to the 63 members NLOHMANN_JSON_PASTE
supports. NLOHMANN_JSON_DERIVED_TYPE_TAG dispatched on the two-token
Type,BaseType prefix plus the member list, running out one slot early:
at 63 members, position 65 landed on the last member name instead of a
sentinel and NLOHMANN_JSON_CAT built an undefined identifier such as
NLOHMANN_JSON_DEFINE_DERIVED_TYPE_INTRUSIVE_m63, with the compiler
reporting "unknown type name 'm1'" once per member and nothing pointing
at an argument-count limit.
That silently reduced all six NLOHMANN_DEFINE_DERIVED_TYPE_* macros from
63 members to 62, contradicting the "up to 63 members" contract in
docs/mkdocs/docs/api/macros/nlohmann_define_derived_type.md.
Drop the leading Type and defer to NLOHMANN_JSON_TYPE_TAG so the tag is
computed from BaseType plus the member list, which fits the available
slots. The zero-own-member derived bodies are therefore selected by tag
1 rather than 2, and the sentinel table for the derived tag is no longer
needed.
Add a regression test at the documented maximum for both the plain and
the derived macros; it fails to compile against the previous dispatch.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Name the zero-member macro bodies by intent, not argument count
The dispatch tag was the literal token 1 or N, pasted onto a macro prefix
to select the zero-member or member-carrying body. For the derived-type
macros that reads wrong: their tag is computed after dropping the leading
Type, so the zero-member body was named _1 while taking two parameters
(Type, BaseType).
Emit EMPTY and MEMBERS instead. The mechanism is unchanged -- the tag is
still a token pasted onto the prefix by NLOHMANN_JSON_CAT -- but the body
names now say what they are rather than encoding an argument count that
only lines up for half of the macros.
Collapse the four duplicated zero-member bodies while here: with no
members there is nothing to default, so each _WITH_DEFAULT_EMPTY body was
a byte-for-byte copy of its plain counterpart. They are now one-line
aliases, leaving a single definition of what an empty object serializes
to per intrusive/non-intrusive and base/derived combination.
No functional change: for both zero-member and member-carrying types the
preprocessed to_json/from_json output is token-for-token identical, and
the arity limits are unchanged (63 members, base and derived).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Document zero-member support in the macro API reference
docs/mkdocs/docs/features/arbitrary_types.md already gained a note, but
the three api/macros pages are where the parameter contract is actually
specified and they still described member as a non-empty list.
State that the list may be empty on each page, and add a note showing
what the zero-member case generates: an empty JSON object for the plain
macros, and base-type-only serialization for the derived ones. Both notes
record that the WITH_NAMES variants do not support this.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Keep user macros named EMPTY or MEMBERS out of the member-count dispatch
The dispatch produced the bare token EMPTY or MEMBERS and pasted it onto
the macro prefix afterwards. In between, the token was rescanned, so a
user macro with either name replaced it: with `#define MEMBERS x` in
scope, even NLOHMANN_DEFINE_TYPE_INTRUSIVE(A, member) -- which compiled
before -- expanded to garbage, and `#define EMPTY` broke the zero-member
form.
Paste the suffix onto the prefix directly in the GET_MACRO slot table
instead. Operands of ## are not macro-expanded, so the selected body name
is formed before any user macro can interfere. NLOHMANN_JSON_TYPE_TAG and
NLOHMANN_JSON_DERIVED_TYPE_TAG become NLOHMANN_JSON_TYPE_BODY and
NLOHMANN_JSON_DERIVED_TYPE_BODY, taking the prefix as their first
argument; NLOHMANN_JSON_CAT is no longer needed. The body macro names are
unchanged, and so is the generated code.
Add a regression test that defines EMPTY and MEMBERS around plain and
derived types, with and without members.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Test for EMPTY and MEMBERS so -Wunused-macros accepts them
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Document the benchmarks, and make them build and compare versions again
The benchmark project hasn't configured since #4793: download_test_data.cmake
compiles cmake/detect_libcpp_version.cpp relative to CMAKE_SOURCE_DIR, which
is tests/benchmarks when that is the top-level project, so try_run fails and
so does `make run_benchmarks`. The path is now relative to the module itself,
which is the same file for the main build.
The Dump benchmark discarded dump()'s result, which is [[nodiscard]] by now;
it warned, and left the optimizer free to shorten the loop. The result is
now kept with benchmark::DoNotOptimize.
A new cache variable, JSON_BENCHMARK_INCLUDE_DIR, names the directory holding
the nlohmann/json.hpp to benchmark (single_include by default, as before),
so the same benchmarks can be built against two versions and compared.
tests/benchmarks/README.md documents what is measured, how to build and run
the benchmarks, how to read the output, and how to compare two versions with
Google Benchmark's compare.py; it recommends doing so by hand before a
release rather than in CI.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Point ci_benchmarks at tests/benchmarks
The target has configured ${PROJECT_SOURCE_DIR}/benchmarks since it was
added in #2561, but the benchmarks live in tests/benchmarks.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Pin Google Benchmark to release 1.9.5
The benchmarks fetched Google Benchmark's main branch, so two builds on
different days could measure with different library code, and CMake 3.30
and later warn that the single-argument FetchContent_Populate() is
deprecated. Fetch the 1.9.5 release archive, verified by its SHA-256,
with FetchContent_MakeAvailable() instead. That needs CMake 3.14; Google
Benchmark itself already needed 3.13.
Its -Werror is switched off, so a newer compiler's new warnings cannot
break the pinned release, and its install rules are no longer added.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Ubuntu, Windows, macOS, and CodeQL already cancel an older run of the same
workflow on the same ref. Check amalgamation, CIFuzz, Dependency Review,
Flawfinder, Semgrep, Scorecard, and the labeler did not, so every push
to a pull request left their earlier runs going. Give them the same
concurrency group. The labeler runs on pull_request_target, where
github.ref is the base branch, so it groups by pull request number.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Answer the OpenSSF Best Practices criteria that asked for policies the
project follows but had not written down:
- SECURITY.md: a first response within 14 days, publishing an advisory
with credit once a fix is released, and that only the latest release
receives security fixes.
- Governance: who has access to the project's resources, how write or
admin access is granted, and how CI secrets are stored and rotated.
- Quality assurance: how dependencies of the build, test, and
documentation tooling are pinned, scanned, and kept free of known
vulnerabilities.
Also update the assurance case, since comparison no longer recurses per
nesting level (#5390), and point the best practices badge and links to
bestpractices.dev under the program's current name.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
#5344 added two lines to unit-deserialization.cpp that clang-tidy
reports: modernize-return-braced-init-list for the remaining() helper
and readability-isolate-declaration for "json j1, j2, j3;".
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Compare values without recursing, and without comparing them twice
Comparing two values compared their containers, which compare their elements,
which brought the comparison back once per nesting level. Two values nested
deeply enough exhausted the call stack and terminated the process with a
segmentation fault - the same bug as #5387, in the last operation that still
had it.
Worse, an ordered comparison took exponentially long in the nesting depth
before C++20. std::vector's operator< is a lexicographical comparison, which
asks whether an element is less than its counterpart and then whether the
counterpart is less than it - two full comparisons of everything below that
element, at every level. Comparing two equal values nested 30 levels deep,
which is nothing unusual, took 3.8 seconds; 40 levels would have taken an
hour, and nothing about the value has to be pathological to get there. C++20
is unaffected: std::lexicographical_compare_three_way asks once.
Compare a value that is nested too deeply to descend into on an explicit
stack instead, in a single pass that yields less, equal, greater or unordered
at once. Equality and the three-way comparison descend as they always did for
the first 128 levels, which nothing measurable costs them; an ordered
comparison no longer descends at all, which is what takes the exponent out of
it. Objects and arrays that are not nested deeply are otherwise compared
exactly as before.
The results are unchanged for every pair of values: 68121 comparisons of a
corpus that covers NaN, discarded values, mixed number types, binary values,
empty containers and both object types are identical to develop, in C++11,
C++17 and C++20, with and without thread_local storage and legacy discarded
comparison. Reproducing that meant reproducing two subtleties: a lexicographic
comparison steps over a pair it cannot order, where a three-way comparison
stops at it, and an object compares its keys with < where its entries are
ordered but with == where they are only checked for equality - not with the
object's own comparator, which for nlohmann::ordered_map tells equality.
Equality needs no ordering, so it no longer asks for any: a key or string type
that can only be compared for equality still works.
Measured (medians of 7 interleaved runs, clang -O3, C++11): comparing two
equal values nested 30 levels deep 3778 ms -> 0.002 ms; ordering flat objects
-33.6%; ordering flat arrays of numbers +27.3%, the one shape that pays for
the single pass; equality unchanged throughout.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Describe comparison in the no-thread-local docs and CI target
Comparing two values now bounds its descent with a thread_local counter
just as copying does, so the JSON_NO_THREAD_LOCAL page, the macro
overview and the ci_test_no_thread_local target cover both rather than
copying alone.
Also record what switching the macro on costs a comparison: on the
benchmark documents, comparing two equal values takes 10% to 90% longer.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Take the descent flag as an argument rather than testing it
MSVC reports the test of a constant as C4127 ("conditional expression is
constant"), which the Windows builds treat as an error: may_descend is
false for operator<, so the operand short-circuits the whole condition.
Passing it to compare_descent_exhausted() puts the test where the value
is an ordinary parameter, and leaves the call sites with no condition of
their own.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Note the comparison fallback in the no-thread-local documentation
The macro page describes what the library defines JSON_NO_THREAD_LOCAL for
by itself in terms of copying alone; comparing falls back the same way.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Parenthesise the reserve() computation in the comparison test
clang-tidy reports the mixed * and + as readability-math-missing-
parentheses, as it does for the identical line in the copy test.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Use the shared descent bookkeeping rather than a second set
Comparing kept a thread_local count, a limit and a guard of its own beside
the ones copying already had, all three the same thing under a different
name. They are gone; the shared count, limit and guard do the work.
The guard grows a second constructor here, because the comparison
operators are written as a macro and a macro cannot use the preprocessor:
it cannot look the count up behind an #ifdef the way copy_structured does,
so the guard looks it up for it. nesting_depth_exhausted() arrives for the
same reason - whether an operator descends at all is a constant at every
call site, and testing it there is what MSVC reports as C4127.
Also say in compare_leaves what happens to a pair that is an array on one
side and an object on the other, since the answer is not obvious from the
code: an operator only descends into two values of the same type, so such
a pair is told apart by its types alone - unequal, and ordered the way the
types are - exactly as it is above the bound.
And record what the explicit stack costs: the comparison operators are
noexcept and the container comparison this replaces allocated nothing, so
running out of memory here ends the process instead of throwing. It takes
a value nested past the bound and an exhausted heap to reach, and the same
comparison used to exhaust the call stack, but it is a new way to fail.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Amalgamate
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* docs: qualify the operator>> stream positioning guarantee
operator>>'s notes state that it leaves the stream positioned right
after the parsed value, so that concatenated JSON values can be read
back to back. That does not hold when the value is a number: a number
is only terminated by the character that follows it, and the lexer's
unget() is simulated (it rewinds only the lexer's own bookkeeping),
so that character stays consumed from the stream.
Document the actual behaviour: the guarantee holds for all value types
except numbers, which must be followed by whitespace. Also qualify the
cross-reference on the JSON Lines page, which repeated the unqualified
claim.
Documentation only; the behaviour itself is tracked in #5340.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* fix: restore the character that terminates a number (#5340)
operator>> is documented to leave the stream positioned right after the
parsed value, so that concatenated JSON values can be read back to back.
That did not hold for numbers: a number is only terminated by the
character following it, and lexer::scan_number() reads that character
and calls unget() -- which is simulated and rewinds only the lexer's own
bookkeeping. input_stream_adapter consumes via sbumpc() with no matching
sungetc(), so the terminating character stayed consumed and the next
extraction started one byte too late ('1true' left the stream at 'rue').
Propagating unget() to the adapter directly does not work: next_unget
makes the following get() replay the cached character, so the terminator
would be delivered twice. Instead, restore the still-pending character
once at the end of a non-strict parse, where the input is handed back to
the caller:
- input_stream_adapter gains unget_character() (sungetc()) and advertises
it via supports_unget, detected the same way as supports_seek.
- lexer::restore_pending_unget() turns a pending simulated unget of a
real (non-EOF) character into a real one and clears next_unget so the
character is not also replayed. It is a no-op for adapters that cannot
unget, and reports failure when sungetc() fails, in which case the
input is left as it was before.
- parser calls it on the three non-strict paths, i.e. for operator>> and
sax_parse(strict = false).
Strict parse()/accept() are unaffected: they require the input to end
after the value, so the character is consumed by the end-of-input check
anyway. Parse error messages and reported positions are unchanged.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* tests: fix CI failures in the #5340 test helpers
Four CI failures, all in the new test code:
- GCC (-Werror=useless-cast): drop the `json(...)` wrapper around
`json::parse(...)`, which already returns a `json`.
- GCC (-Werror=unused-result): assign the discarded `json::parse()`
result to a dummy, the idiom used elsewhere in the test suite, and
catch `json::parse_error&` for consistency.
- clang-tidy (google-default-arguments): remove the default argument
from the `pbackfail()` override; `sungetc()` supplies the base
declaration's default.
- MSVC (bad allocation): `no_putback_streambuf::underflow()` set a
one-character get area without advancing `m_pos`, so an implementation
whose `istream::get` peeks before it bumps re-read the same character
forever. Keep no get area at all: `underflow()` peeks, `uflow()`
consumes, and `sungetc()` still always lands in `pbackfail()`, which
is what the test needs.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* fix: leave the character that terminates a number in the input
Read the character following a number without consuming it, instead of
consuming it and putting it back. input_stream_adapter now peeks with
sgetc() and only steps over the character when the next one is requested
or when the adapter is destroyed, so releasing it cannot fail - no
putback position is required from the streambuf.
Suggested by gregmarr in #5344.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* docs: match the version history wording to the peek-based fix
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* docs: drop the whitespace-separator caveat from the parsing pages
The caveat added in #5343 describes the behavior this branch fixes: a
number no longer consumes the character that terminates it, so
concatenated values need no separator.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* refactor: split the strict and non-strict paths in parser
Folding the release_lookahead() call into the existing strict check left
the "in strict mode" comment on an else-if branch, and made the strict
condition in sax_parse() redundant with the branch it followed.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Put the stream position fix behind JSON_PRECISE_STREAM_POSITION
Leaving the character that terminates a number in the stream is observable:
reading "1,2,3" with repeated operator>> works today only because the comma
after each number is swallowed, and std::getline after a number skips the
line break. Both break with the fix, so make it opt-in for 3.x, as suggested
by @gregmarr in the review.
- JSON_PRECISE_STREAM_POSITION (default 0) selects the peek-based
input_stream_adapter. Without it, the adapter is the consuming one from
develop and has no supports_lookahead, so lexer::release_lookahead() and
the parser's calls to it compile to nothing.
- The macro changes input_stream_adapter's layout and member functions, so
it gets the ABI tag _psp, after _bics. The ABI config tests, the natvis
generator, and nlohmann_json.natvis (regenerated) know the tag.
- The tests for the fix move to unit-precise-stream-position.cpp, which
defines the macro itself and runs in every build, and gain the two cases
above. unit-deserialization.cpp pins the default behavior instead.
- The docs describe the default behavior again and point to the new macro
page; version history says "added in 3.13.0, planned default in 4.0.0".
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The windows-11-arm runner now ships MSVC 19.51, which reports doctest's
forward declaration of std::tuple as C5285 ("cannot declare a
specialization for 'std::tuple'"). With /WX this breaks the msvc-arm64
job on develop and on every open pull request. Disable the warning for
the test targets, like the other MSVC warnings already disabled there.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Keep JSON_DIAGNOSTICS parent pointers of ordered_json members after erase() and update()
ordered_json stores its members in a vector, and two operations moved
members without restoring their parent pointers afterwards:
- ordered_map::erase() re-constructs every member after the erased one in
place. The basic_json move constructor leaves m_parent at nullptr, and
none of the object branches of basic_json::erase() (by key, iterator, or
iterator range) called set_parents(). This also affected merge_patch()
with a null member and patch() with a remove operation.
- update() only set the parent pointer of the inserted member. Adding a key
can reallocate the vector, which copies all other members and leaves
their m_parent at nullptr. The set_parents() call added for #4813 only
repaired this for the nested object of a merge, not for the target.
The next assert_invariant() on such an object (for instance, when copying
it) aborted, and diagnostic messages lost the path prefix above the moved
member. std::map-based json was not affected, because its nodes do not
move.
Erasing from an ordered_map object now calls set_parents(), and update()
uses set_parent(), which already refreshes all members for vector-based
objects. This makes the #4813 workaround redundant.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Account for JSON_DIAGNOSTIC_POSITIONS in the ordered_json parent-pointer test
The merge_patch() case parses its input, so with JSON_DIAGNOSTIC_POSITIONS
the exception message also carries the byte range of the parsed value.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Silence clang-tidy for the intentional copy in the ordered_json parent-pointer test
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Keep parent pointers when update() merges past its descent bound
The iterative path of update() only set the parent pointer of the member
it inserted, like the recursive one did before. It now uses set_parent()
too, so ordered_json members that move when a nested object grows keep
their parents, and the set_parents() calls that patched this up after
each nested merge are gone.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Check the fuzzers' UBJSON/BJData round-trip invariants in the unit tests
The strongest correctness checks for the UBJSON and BJData writers lived
only in the OSS-Fuzz drivers: anything from_ubjson()/from_bjdata()
returns must serialize with every option combination, parse back, and
re-serialize stably. Those checks only run at OSS-Fuzz, so regressions
surfaced days later as external reports - the same BJData assert pair
was reported five times over three years, and #5494's harness change
was followed by OSS-Fuzz 563659413 within a day.
Add "UBJSON round-trip invariants" and "BJData round-trip invariants"
test cases that run the drivers' checks on a fixed, deterministic corpus
(tests/src/round_trip_corpus.hpp): integer and float boundaries,
non-finite numbers, strings, binary values, optimized containers, deep
nesting, the JData annotated-array matrix, and seeded random containers.
They also check two properties the drivers do not: the first round trip
preserves the value, and re-serializing reproduces the exact bytes. For
BJData both exclude values containing a binary value, which is read back
as an array of integers unless it was written as a Draft 3 optimized
binary array; this carve-out is now documented in bjdata.md. Run against
the headers before #5542, the BJData test fails, including on the shape
from OSS-Fuzz 563659413.
Also document how OSS-Fuzz reports are handled (reference them as
"OSS-Fuzz: <id>", turn the reproducer into a unit test, keep drivers and
unit tests in sync) in tests/fuzzing.md, and link it from the PR
template and the quality assurance page.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add the OSS-Fuzz reproducers for 474400817 and 474480402 as unit tests
Following the convention added to tests/fuzzing.md, the reproducers of
the two BJData fuzzer asserts tracked since January are now unit tests:
- 474400817 (assert(false)): an empty object _ArraySize_ was written as
the ND-array header length, which from_bjdata() could not read back.
Fixed by #5455.
- 474480402 (to_bjdata(j2, false, false) == vec2): a one-byte Draft 3
binary array is written in Draft 2 mode as a uint8 array and then
re-serialized with the int8 marker. This is the documented exception to
byte stability, not a library bug; OSS-Fuzz closed it after #5494
relaxed the harness to value stability. The test pins the exact bytes
so the exception stays deliberate.
The 563659413 reproducer is already a unit test (#5542). A comment also
ties the existing UBJSON excessive-count test to the timeout OSS-Fuzz
reported for that shape (testcase 6347769435193344).
OSS-Fuzz: 474400817
OSS-Fuzz: 474480402
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix GCC -Weffc++ and -Wuseless-cast warnings in the round-trip corpus
Initialize the atoms in the member initialization list, and drop the cast of
the generator's result, which already is std::size_t on 64-bit Linux.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Describe the threat model, the trust boundaries, the secure-design
argument, and how common weaknesses are countered, with links to the
quality assurance page as evidence.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Describe what the project will and will not do over the next year,
and point to issue #3453 for the open question of a 4.0 release.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Complete the architecture documentation page
Replace the placeholder bullets and TODOs with a description of the
component pipeline (with a diagram), the source layout, the template
parameters, the value storage (now in struct data), the input and
output adapters, and the SAX interface.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Link sources and basic_json, document full input adapter interface
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Align the default column of the template parameter table
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The windows-11-arm runner image moved to windows-11-vs2026-arm64, which no
longer ships Visual Studio 2022, so the msvc-arm64 job failed at configure
time. Use the "Visual Studio 18 2026" generator like the msvc2026 job.
clang-tidy's modernize-raw-string-literal check flagged two string literals
in the nesting tests added by #5546 and #5547.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Merge deeply nested objects without recursing per nesting level
merge_patch() and update(j, true) merged a nested object by calling
themselves on it, once per nesting level. A value nested deeply enough -
50,000 levels of objects on an 8 MiB stack - exhausted the call stack
and terminated the process, although parse() accepts such values without
complaint.
Bound the descent the same way dump() does. The recursion now carries
the nesting level, and once merge_depth_limit() (128) levels have been
entered, update_members_iteratively() and merge_patch_iteratively()
finish the merge on an explicit stack. They still merge a nested object
completely before the next member, and in the same order, so the results,
including the parents JSON_DIAGNOSTICS reports paths from, are unchanged.
Values nested less deeply than the bound run the same code as before, so
the common case does not pay for the stack: merging only on it cost
10-14% in a first version.
The public signatures are unchanged. The recursive worker behind
merge_patch() has its own name rather than being a private overload, so
that &basic_json::merge_patch stays unambiguous.
Tests check every depth up to 300 against recursive reference
implementations of both operations, check the diagnostic paths past the
bound, and merge objects nested 100,000 levels deep.
Fixes#5545 for update(j, true), and #5393 for merge_patch().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Use the shared recursion limit in update() and merge_patch()
merge_depth_limit() is gone in favor of detail::recursion_depth_limit().
The two identical function-local frame structs become one member struct,
merge_frame, with a constructor, so both loops emplace_back() their
frames. merge_patch_iteratively() copies the frame it works on out of the
stack and changes it only through stack.back().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Build the update()/merge_patch() diagnostics test values instead of parsing them
Parsed values carry byte positions under JSON_DIAGNOSTIC_POSITIONS, which
the expected messages do not include.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Hash deeply nested values without recursing per nesting level
std::hash<basic_json> hashed an array or object by hashing each element,
which called detail::hash again once per nesting level. A value nested
deeply enough - 50,000 levels of objects on an 8 MiB stack - exhausted
the call stack and terminated the process. parse() accepts such values
without complaint, since the parser is iterative, and a parsed value is
hashed wherever it is used as a key in an unordered container.
Bound the descent the same way dump() does: detail::hash takes the
nesting level, and once hash_depth_limit() (128) levels have been entered,
hash_iteratively() hashes what is left on an explicit stack. It combines
the seeds in exactly the same order, so hash values are unchanged. A value
nested less deeply than the bound is hashed by the same code as before,
without allocating, and is as fast as before.
Tests check that every depth up to twice the bound hashes exactly like
the recursive definition of the hash, and that values nested 100,000
levels deep hash without crashing.
Fixes#5545 for std::hash.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Declare hash_frame's constructor noexcept
GCC's -Wnoexcept (an error in CI) flags the emplace_back() into the
hash stack under C++26: the constructor cannot throw, since cbegin() is
noexcept, but it did not say so. dump_frame's constructor is noexcept
for the same reason.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Share one recursion depth limit, and copy the hash frame out of the stack
dump() and hash() each defined their own limit on how many nesting levels
they recurse into, and the operations still to come would have added more,
free to diverge over time. They now all use detail::recursion_depth_limit(),
in a header of its own; serializer::dump_depth_limit() and
hash_depth_limit() are gone.
hash_iteratively() now copies the frame it works on out of the stack and
changes the frame only through stack.back(), so nothing can refer into
the stack after entering an element has grown it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Parenthesize multiplications in the hash test for clang-tidy
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add missing headers to BUILD.bazel and make its generator reproduce it
The "json" cc_library did not list three headers that the library
includes:
- detail/meta/logic.hpp (added in #5016, included by from_json.hpp)
- detail/input/number_parse.hpp (added in #5283, included by lexer.hpp)
- detail/input/string_scan.hpp (added in #5283, included by lexer.hpp
and serializer.hpp)
Bazel's sandbox only exposes declared headers, so any target depending
on @nlohmann_json//:json and including <nlohmann/json.hpp> failed with
"'nlohmann/detail/meta/logic.hpp' file not found".
The file could not simply be regenerated, because the generator behind
"make BUILD.bazel" was stale: it wrote only the "json" cc_library and
dropped the load() statements, the license block, and the
"singleheader-json" target that were added by hand in #4584. The
generator now emits the complete file, so its output differs from the
previous BUILD.bazel only by the three headers. It also resolves the
glob against the project root instead of the working directory and
sorts the list explicitly.
"make BUILD.bazel" is now phony: in a fresh checkout, BUILD.bazel is
not older than the headers, so make considered it up to date, and a
removed header would never trigger a rebuild. "make check-amalgamation"
also checks that BUILD.bazel is up to date.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Check in CI that BUILD.bazel is up to date
The "Check amalgamation" workflow now also regenerates BUILD.bazel, so a
pull request that adds, renames, or removes a header without updating
the Bazel header list fails, and the attached amalgamation.patch
contains the fix. The failure comment and the contribution guidelines
mention the new check, and the comment now links to the existing
"Amalgamate the source code" section instead of the "Files to change"
anchor that was removed in #4560.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Test JSON_BRACE_INIT_COPY_SEMANTICS for real, and fix one-element tuples under it
The opt-in JSON_BRACE_INIT_COPY_SEMANTICS was never exercised by CI:
- Its only test, in unit-regression3.cpp, was guarded by
`#if defined(JSON_BRACE_INIT_COPY_SEMANTICS)` after the #include. The
header #undefs the macro unconditionally in macro_unscope.hpp, so the
guard was always false and the test compiled to nothing, whatever -D
flag was passed.
- The ci_test_brace_init_copy_semantics target that passes the flag was
not named by any workflow.
Move the test into its own translation unit that defines the macro before
including the header, as unit-diagnostics.cpp does for JSON_DIAGNOSTICS.
It now runs in every CI job and for every standard. Remove the unused
target: it ran the whole suite with the macro, and that suite deliberately
relies on default brace-init semantics in about 90 places
(e.g. `json({1})` meaning `[1]`), so it could never pass.
Running the whole suite with the macro did find one library bug:
to_json for std::tuple builds `j = { std::get<Idx>(t)... }`, so with copy
semantics a one-element tuple became its element. `json(std::tuple<int>{5})`
was `5` instead of `[5]`, and `get<std::tuple<int>>()` threw type_error.302
on the result. Under the macro, a one-element tuple now builds exactly what
the default deduction builds. Without the macro nothing changes.
The new tests also pin that the library's other conversions produce the
same values with and without the macro. The macro page now says that the
macro affects every single-element list (`json j = {1}` is `1`), and that
all translation units must agree on it, since it has no ABI tag.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Make JSON_BRACE_INIT_COPY_SEMANTICS part of the ABI tag
The macro changes the body of the initializer-list constructor and adds a
to_json_tuple_impl overload, both with the same mangled names in either
mode, so mixing translation units silently picked one definition. Encode
it in the inline namespace as `_bics`, as JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON
does with `_ldvcmp`. The macro is new in the unreleased 3.13.0, so no
existing namespace name changes.
- Move the macro's default into abi_macros.hpp so json_fwd.hpp computes
the same namespace, and keep it defined under JSON_TEST_KEEP_MACROS.
- Check the tag in the ABI config tests and in the unit test.
- List `_bics` (and the missing `_dp`) in the namespace docs and in the
natvis generator; regenerate nlohmann_json.natvis.
- Replace the "define it consistently" warning with an ABI note.
Suggested by @gregmarr in the review of #5544.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix the cppcheck, clang-tidy and legacy-comparison CI failures
- to_json_tuple_impl() moved the element in both branches of a ternary;
only one runs, but cppcheck reported accessMoved. Use if/else.
- The ABI tag test looked for "json_abi_bics", which misses when another
tag comes first, as in json_abi_ldvcmp_bics; look for "_bics".
- readability-qualified-auto in the items() test.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
write_ubjson() built a std::vector of the eight markers BJData forbids as
the type of an optimized container - one heap allocation plus a linear
search for every array and object it wrote with use_type, even for plain
UBJSON output, where the list isn't consulted. The list was also spelled
out twice. A constexpr helper, is_bjdata_excluded_type_marker(), replaces
both.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- "aspect: binary formats" for changes to the binary reader or writer,
their tests, fuzzers and docs, or with a binary format in the title;
- "python" for Python sources and pip requirements files, matching the
label Dependabot sets on its pip updates, so it is never removed there;
- "CI" also for changes to the Dependabot and labeler configurations.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The fuzzer drivers check their round trips with assert(), which NDEBUG
compiles away. The OSS-Fuzz build keeps assertions on today, but nothing
pins that: a build change that adds NDEBUG would silently turn every
round-trip check into a mere "does not crash" check. Each driver now
stops the build with an #error instead, and includes <cassert> itself
rather than relying on json.hpp.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The community-maintained clang-tidy check
modernize-nlohmann-json-explicit-conversions rewrites implicit
conversions into explicit get<T>() calls, which is exactly the
preparation the docs ask for ahead of implicit conversions being
switched off by default. Mention it on the JSON_USE_IMPLICIT_CONVERSIONS
page and in the migration guide, as promised in discussion #4610.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Without a bind address in serve_header.yml, the server listened on all
interfaces, so any machine on the network could fetch the header and
trigger make runs in the working trees. It now listens on localhost
unless configured otherwise; bind: null restores the old behavior.
The header was also sent with Access-Control-Allow-Origin: *, letting
any web page read it. CORS is only needed because Compiler Explorer
downloads #include <https://...> headers in the browser, so the header
now goes only to https://godbolt.org and https://compiler-explorer.com,
configurable with cors_origins.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
write_bjdata_ndarray() encoded a JData-annotated object as a BJData
ND-array whenever its dimensions' product matched _ArrayData_.size(),
which lost information in two ways:
- _ArrayData_ was never required to be an array. null has size 0, any
other scalar has size 1, and iterating an object visits its values, so
e.g. {"_ArraySize_":[1],"_ArrayData_":5} was written as the array [5],
and an object _ArrayData_ came back as an array.
- The reader only restores an annotated object from an ND-array with at
least two non-zero dimensions that is not a 1xN row vector; an empty,
1-D, row-vector, or zero-sized shape is read back as a plain array. The
writer nonetheless emitted ND-array headers for these shapes, so the
annotation was silently dropped.
OSS-Fuzz issue 563659413 hit this in parse_bjdata_fuzzer: an empty binary
_ArraySize_ is written as a plain object and read back as an empty array,
after which {"_ArrayType_":"int16","_ArraySize_":[],"_ArrayData_":null}
was encoded as the ND-array header "[$I#[]" and re-read as [], failing the
harness's value-stability check.
Such objects now fall back to a plain object encoding, which round-trips.
Genuine ND-arrays (two or more positive dimensions, not a 1xN row vector)
are encoded exactly as before. Existing fallback tests that used 1-D
shapes are moved to 2-D shapes so they keep exercising the check they
were written for, and the BJData documentation is updated.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Match ABI tag order in namespace tests to abi_macros.hpp
NLOHMANN_JSON_ABI_TAGS concatenates the tags as _diag, _ldvcmp, _dp,
but the default and noversion ABI tests expected _diag, _dp, _ldvcmp.
The tests therefore failed whenever both JSON_DIAGNOSTIC_POSITIONS and
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON were enabled, a combination
CI never exercises. Reorder the expectations to match the header.
Also document the _dp tag in the namespace feature page, which listed
only _diag and _ldvcmp.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Test the ABI namespace with all ABI tags enabled
Build the default and noversion ABI config tests a second time with
JSON_DIAGNOSTICS, JSON_DIAGNOSTIC_POSITIONS and
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON all set, so the expected tag
order is checked on every test run instead of depending on which CMake
options a CI job happens to enable.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Bound the descent of the copy constructor
basic_json's copy constructor copied objects and arrays by handing the
container to its own copy constructor, which copy-constructs every element
and so reaches this constructor again, once per nesting level. A value
nested deeply enough exhausted the call stack and terminated the process
with a segmentation fault - no exception, nothing the caller could catch.
Parsing such a value works, as the parser is iterative, and so does
destroying one, as #1436 made destruction iterative.
Bound how far the copy descends rather than take the call stack away from
it. The first levels are copied exactly as they were - the containers copy
their own elements, which is by far the fastest way to fill them - and only
once the copy has descended 128 levels is the value below it finished
without the call stack, through an explicit worklist. Copying can therefore
no longer exhaust the stack, however deeply a value is nested, while a value
nested less deeply than the bound - all but a vanishing minority - is copied
by the very same code as before and pays only for one counter.
That counter lives in thread_local storage, as one shared between threads
would be raced. JSON_NO_THREAD_LOCAL switches it off for toolchains without
thread_local; copying then goes through the worklist right away, which
yields the same values but is measurably slower.
The deferred values are completed before the copy they belong to returns, so
a value copied while another copy is going on - by a custom base class, say -
is unaffected by the copy it is nested in.
operator= takes its argument by value, so copy assignment is fixed as well.
Copying is as fast as it was, within measurement noise (medians of 9
interleaved runs, clang -O3): -1.3% for an array of strings, +0.0% for a
flat object, +0.1% for a flat array of numbers, +0.3% for nested arrays,
+0.6% for nested objects and +1.2% for a twitter-like document. Copying a
three-key object costs about ten nanoseconds more, the counter. Deferring
every level instead, rather than only those below the bound, measured
between 3% and 9% slower depending on the shape of the value.
This fixes#5387 for the copy constructor. dump() is still recursive.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Test the copy constructor's iterative path in CI
The copy constructor descends into 128 levels before it finishes a value
without the call stack, so the iterative path is otherwise only reached
by the few tests that nest deeper than that.
JSON_NO_THREAD_LOCAL switches the descent off, which sends every value
down that path. Running the whole test suite that way covers it with
every object type, string type, allocator, and base class the suite
already exercises. The new ci_test_no_thread_local target does that; the
macro had no build coverage at all before.
Copying a nested value also has to carry over what the element-wise copy
constructor would have copied: the parents that JSON_DIAGNOSTICS relies
on, and the positions that JSON_DIAGNOSTIC_POSITIONS reports. Both are
now checked on either side of the descent bound, for objects and arrays.
Neither was tested before, and dropping either one makes the new tests
fail.
Also quantify what JSON_NO_THREAD_LOCAL costs a copy instead of calling
it "measurably slower".
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Split the regression tests so that they keep linking
Linking test-regression2 fails with "relocation truncated to fit:
IMAGE_REL_AMD64_REL32 against `.rdata'" once its object grows past what
the MinGW linker copes with, and the copy constructor's helpers push it
over: the object grows by 6.3%, from 4,654,128 to 4,944,920 bytes at -O0,
and develop links at the smaller of the two.
Building the tests optimized shrinks the object enough to link, but the
binaries clang 11.0.1 and clang 18.1.8 then produce crash before doctest
prints its first line - 39 of 102 tests on clang 18 - so the objects have
to become smaller rather than denser.
Moving the test cases that follow "regression tests 2" into a file of
their own brings that object to 4,687,888 bytes, which is 0.7% above the
size that links today rather than 6.3%. Both files still build for C++11,
C++17 and C++20, and run the same 9 test cases and 135 assertions as
before, now spread over two binaries.
New regression tests belong in unit-regression3.cpp from here on, which
is what CONTRIBUTING.md now says.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Do not use thread_local storage with Clang targeting MinGW
Every test that copies a value segfaults there - 42 of 105 on clang
11.0.1, 39 of 102 on clang 18.1.8 - while the same tests pass with GCC
targeting MinGW, with Clang targeting MSVC, and with every other
toolchain the library is tested on. The counter that bounds the copy
constructor's descent is the library's first use of thread_local, so
that job had never exercised it before.
JSON_NO_THREAD_LOCAL already covers toolchains without thread_local
storage, and copying yields the same values with it, only more slowly.
Define it for this one automatically.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Balance the warning suppression the split separated
unit-regression2.cpp opens a DOCTEST_CLANG_SUPPRESS_WARNING_PUSH block at
the top and closed it at the very bottom, which the split moved into
unit-regression3.cpp: one file was left with a push and no pop, the other
with a pop and no push, which clang reports as an error.
Give each file the pair it needs.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Check both shapes without a C-style array
clang-tidy rejects the array the two shapes were iterated over
(cppcoreguidelines-avoid-c-arrays). The array only existed because astyle
reformats a range-for over a braced initializer list into something
unreadable; naming the two cases avoids both.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Split the regression tests far enough to leave room
The first split left unit-regression2.cpp 0.7% below the size develop
links at, which the comparison change in the follow-up immediately used
up: the MinGW linker fails on test-regression2_cpp20 again, naming
copy_shallow and to_partial_ordering among the relocations it cannot fit.
Move the sections from "issue #2067" on, and the helper types they use,
so that the file stops being the one that decides whether the tests can
be linked at all. At -O0 and C++20, unit-regression2.cpp is now 2,964,944
bytes against develop's 4,708,248, and 3,070,568 bytes with the follow-up
applied - roughly a third smaller either way, rather than a fraction of a
percent larger.
The 135 assertions are the same ones as before, now spread over three
test cases in two files.
Also silence the clang-tidy findings the deep-nesting tests draw: the
copies they make are what is being tested, and the reserve() computation
gets its parentheses.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move the #4804 alias to the file that uses it
The split left the json_4804 alias behind in unit-regression2.cpp while
the test case that uses it went to unit-regression3.cpp, which does not
build for C++17 and C++20 as a result.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Include <span> where the split moved its only use
The #2546 test case guards itself with __has_include(<span>), but the
include itself sat in unit-regression2.cpp's preamble and stayed behind,
so the section compiled without a declaration wherever the guard passed -
which nvhpc reported and libc++ builds do not, as they skip the section
altogether.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Keep the descent bookkeeping in one place
Copying carried a depth count, a depth limit and a guard of its own, and
the comparison in the follow-up added a second set beside them. Neither
operation needs its own: they are never nested inside one another by the
library - copying a value does not compare one, and comparing two values
does not copy them - and where user code nests them anyway, sharing the
count only ends a descent sooner than it had to.
So there is now one nesting_depth(), one nesting_depth_limit() and one
nesting_depth_guard, which the follow-up uses instead of adding its own.
Inverting the test in copy_structured leaves the too-deep case and the
no-thread-local case as the same code.
The guard takes the count rather than looking it up, because the caller
has looked it up already to test it against the limit, and reaching
thread-local storage twice on the path that is taken almost every time is
worth avoiding.
The switch that copies the value of anything that is not an object or an
array was written twice - once in the copy constructor, once in
copy_shallow - so that adding a value_t meant editing both, and missing
one would have been silent. It is copy_leaf_value now, and inlined: both
callers have already sorted the containers out, and folding that test into
the switch is what keeps a value made mostly of numbers copying as fast as
it did.
Copying canada.json, citm_catalog.json and twitter.json is within 0.6% of
what it was before, measured as a paired ratio over 18 interleaved rounds
against a run-to-run spread of 0.3%.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Check that an abandoned copy can still be destroyed
Copying a value without the call stack builds the copy from the top down,
and every value whose own copy has not been made yet stays a null value
until it is. That is what lets a copy be abandoned half-built: the
destructor finds nothing but complete values and null ones.
Nothing tested it. Failing an allocation part-way through a copy of a
deeply nested value does, with the allocator the file already has for
exactly this kind of test.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Name the test's locals so Flawfinder stops matching them
The code scanning job reports CWE-362 - "check when opening files" - for
a test that opens no files: Flawfinder matched a local variable called
open. Rename it and its partner.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Keep the descent guard's bookkeeping self-contained
nesting_depth_limit() and nesting_depth_guard were only used inside
the JSON_NO_THREAD_LOCAL-guarded branch of copy_structured(), but were
defined unconditionally. Move them inside the #ifndef, and have the
guard look up the depth and test it against the limit itself (via
okay()) instead of making the caller do it - the caller no longer
needs to touch nesting_depth() at all. Also shrink the thread-local
counter to std::uint8_t, matching what its own doc comment already
argued.
Addresses gregmarr's review comments on #5389.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Make nesting_depth_guard usable regardless of JSON_NO_THREAD_LOCAL
nesting_depth_limit() and nesting_depth() stay behind #ifndef
JSON_NO_THREAD_LOCAL, since a descent cannot be bounded without a
per-thread count. But the guard itself now always exists, becoming a
no-op that is never okay() under that macro - the same way the bound
is already reached on every call without one. copy_structured() no
longer needs to know which case it is in.
This is what lets #5390 reuse the guard for comparison, which cannot
test JSON_NO_THREAD_LOCAL where the macro-based operators use it: the
guard now carries that distinction itself instead of requiring every
caller to.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Silence VS2015's C4503 for the custom-base-class test
The deep-copy support added for #5387 lengthened the mangled name of
std::allocator_traits<...>::construct for the test's map type past
VS2015's limit, which /WX turns into a build failure even though the
name is only used for (now-truncated) debug info.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Remove dead unused-parameter casts from copy_metadata()
@gregmarr asked whether the static_cast<void> pair in the
JSON_DIAGNOSTIC_POSITIONS-off branch was needed for an empty
json_base_class_t. It isn't: src and dst are already referenced
unconditionally by the base-class copy above, so no -Wunused-parameter
warning fires either way (checked with -Wall -Wextra
-Wunused-parameter, JSON_DIAGNOSTIC_POSITIONS 0 and 1).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix CI: build custom array types without a fill constructor, re-amalgamate
copy_array_level() built the destination array with the fill
constructor array_t(count, value), which is not part of the array
container interface the library otherwise assumes (e.g. custom
ArrayTypes that only provide a default and an iterator-pair
constructor, as covered by unit-custom-array-type.cpp). Default-
construct the array and resize() it instead, matching how the rest
of the codebase already grows array_t.
Also re-run the amalgamation, which had fallen out of sync with
include/nlohmann/json.hpp.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* docs: document that a NUL byte in the input is treated as end of input
A NUL byte anywhere in the input - trailing, or embedded ahead of more
otherwise well-formed JSON - is currently treated the same as genuine
end of input, so parsing silently stops there instead of raising the
parse_error.101 any other unexpected byte triggers. This mirrors the
NUL-terminated-C-string convention already used when no explicit input
length is given (json::parse(const char*) already stops at strlen()),
just applied uniformly rather than only when a length is genuinely
unavailable.
This behavior predates this change and is not being altered here -
changing it would be an observable, backwards-incompatible behavior
change for any caller that (knowingly or not) depends on it, which is
not something to do silently in a patch. Documenting the current,
verified behavior as a new FAQ entry instead, so it's an intentional
and discoverable part of the contract rather than a surprise.
Fixes#5530.
Signed-off-by: Niels Lohmann <niels.lohmann@gmail.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4RQ1Ahan5YAGbnAQGjZTY
* Add JSON_STRICT_NUL_HANDLING opt-in macro for issue #5530
A NUL byte anywhere in the input is currently treated the same as real
end of input, rather than raising parse_error.101 like any other
unexpected byte (documented in the previous commit's FAQ entry). A full
unconditional fix was tried in PR #5532 but rejected as too risky to
ship by default: any caller could depend on the current behavior, even
unknowingly (e.g. a zero-padded buffer). On PR #5534, gregmarr proposed
a compile-time opt-in flag instead, and the maintainer agreed, wanting
it available now and defaulting to the corrected behavior in 4.0.0.
This mirrors the existing JSON_BRACE_INIT_COPY_SEMANTICS precedent as
closely as sensible:
- JSON_STRICT_NUL_HANDLING defaults to 0 (off); the three lexer sites
that treat '\0' as EOF/comment-terminator are gated with
`#if !JSON_STRICT_NUL_HANDLING` so the default-off behavior is
byte-for-byte identical to today's.
- input_adapters.hpp's `T (&array)[N]` overload additionally trims a
single trailing '\0' from a `char` array (e.g. a string literal like
`json::parse("123")`) when the macro is on, so that case keeps
working; every other element type (unsigned char, std::uint8_t, ...)
always keeps its full extent. This intentionally does *not* reuse the
existing strlen()-based pointer overload via SFINAE-excluding `char`
from the array overload, as originally sketched for this change: that
approach is ambiguous against the newer generic container overload
added since PR #5532, and even where it compiles, strlen()-scanning a
`char` array that is not NUL-terminated within its bounds reads past
the end of the array (confirmed with AddressSanitizer). Trimming only
a single trailing byte, without scanning, avoids both problems.
- Documented via docs/mkdocs/docs/api/macros/json_strict_nul_handling.md,
linked from the macros index/nav/features page, the FAQ entry, and
the parse/accept/operator>> reference pages.
- Tested in unit-class_parser.cpp and unit-deserialization.cpp, default
state unguarded and opt-in state guarded. Since the library itself
#undefs the macro at the end of json.hpp (as JSON_BRACE_INIT_COPY_SEMANTICS
already does), a plain `#if defined(JSON_STRICT_NUL_HANDLING)` guard
after the include never actually triggers; the tests instead capture
the command-line value into a test-local macro before including the
header. A few pre-existing fixtures elsewhere (std::array<uint8_t, N>
sized one larger than their literal, relying on value-initialization
to silently add a trailing zero byte) needed the same one-byte
adjustment to keep passing under the opt-in behavior.
Unlike the precedent, this adds a proper `JSON_StrictNulHandling` CMake
option (rather than a raw -DCMAKE_CXX_FLAGS injection) and wires its
ci_test_strict_nul_handling target into the ci_cmake_options job matrix
in .github/workflows/ubuntu.yml, so the opt-in build is actually
exercised in CI -- closing the one gap in the precedent's own CI setup
(ci_test_brace_init_copy_semantics is defined but never referenced by
any workflow, so it has never actually run).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Clarify where JSON_STRICT_NUL_HANDLING does not reject NUL bytes
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <niels.lohmann@gmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Add 60 contributors whose work was not yet credited and update seven
links that pointed to renamed or reassigned GitHub accounts.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Devirtualize binary_writer via a value-type output sink
to_cbor/to_msgpack/to_ubjson/to_bjdata/to_bson wrote every byte through
output_adapter_t, a shared_ptr<output_adapter_protocol> whose
write_character/write_characters are virtual. Unlike the lexer (templated
on a concrete InputAdapterType), the binary writer never got that
treatment, so binary output paid a vtable lookup per byte and a
make_shared per call.
Template binary_writer on an OutputSinkType and give it two concrete,
non-virtual sinks:
- output_vector_sink: appends straight into a std::vector (push_back /
insert), used by the vector-returning to_* convenience functions. No
vtable, no shared_ptr; the writes inline.
- output_adapter_sink: forwards to a type-erased output_adapter_t, so the
existing to_*(j, output_adapter) overloads (streams, strings, custom
adapters) keep working exactly as before -- one virtual call each,
unchanged.
binary_writer keeps a convenience constructor taking output_adapter_t
(building the default output_adapter_sink), so the adapter overloads are
untouched; only the convenience functions switch to the vector sink. The
friend declaration and the basic_json binary_writer alias gain the new
(defaulted) template parameter.
Output is byte-for-byte identical: verified across ~3000 randomized
values plus curated edge cases (all scalar widths, strings with invalid
UTF-8, binary, nested arrays/objects) for CBOR, MessagePack, UBJSON (both
size/type settings), BJData, and BSON, plus the output_adapter path, in
C++11/17/20. Warning-clean under clang -Weverything and the gcc pedantic
set; clang-tidy clean on the changed headers; make check-amalgamation
clean.
Throughput (g++ -O3, vs develop): scalar-dense binary output such as
integer arrays ~1.4x; many small to_cbor calls ~1.04x (DOM traversal
bound); string/blob-heavy output unchanged (already bulk-bound). No
workload regressed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix CI failures from binary_writer output-sink change
Four CI jobs failed on the initial commit; all are addressed here without
changing any output (binary encodings remain byte-for-byte identical to
develop across the differential corpus):
1. ci_test_gcc / cuda (-Werror=duplicated-branches): for number_float_t ==
float, static_cast<float>(n) is the identity, so write_compact_float's
two branches are intentionally identical. Once the concrete vector sink
is inlined, GCC constant-folds and diagnoses this (the type-erased path
hid it behind a non-inlined virtual call). Silence -Wduplicated-branches
for GCC (clang has no such warning) alongside the existing -Wfloat-equal
pragma.
2. ci_static_analysis_clang (UBSan nonnull-attribute): binary_writer passes
a null pointer with length 0 for empty strings/binary. output_vector_sink
/ output_adapter_sink declared write_characters JSON_HEDLEY_NON_NULL, so
the sanitizer flagged the (harmless) zero-length call once the sink was
called directly rather than through the attribute-free virtual base. Drop
the attribute from both sinks, matching the pre-existing behavior.
3. ci_cpplint (build/include_what_you_use): output_adapter_sink uses
std::move; add #include <utility>.
4. ci_cuda_example (nvcc 11.8): NVCC's front end rejects the default
template argument on the binary_writer alias template. Revert the alias
to its original single-parameter form (relying on binary_writer's own
defaulted OutputSinkType) and spell out the full type in the vector-sink
convenience functions.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Encode big-endian numbers with a byte swap instead of std::reverse
write_number() reordered multi-byte numbers for the big-endian formats
(CBOR/MessagePack/UBJSON) with std::reverse over the byte array. GCC
lowered only some sizes to a bswap; clang kept a scalar byte shuffle
(0 bswap instructions in the CBOR number path). Replace the reverse with
size-dispatched __builtin_bswap16/32/64 helpers (portable shift fallback
for other compilers; std::reverse retained for exotic sizes such as a
long double number_float_t).
Codegen: the CBOR number path now emits bswap on both compilers
(gcc 2 -> 16, clang 0 -> 4). Output is byte-for-byte identical to the
previous implementation across the binary differential corpus.
Throughput (isolated vs the std::reverse version, best of 9):
CBOR int64 array gcc +7% clang +10%
CBOR uint16 array gcc +27% clang flat
Modest but consistent on number-dense encodings; negligible on
string/blob-heavy output, as expected.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Reserve output capacity up front for binary serialization
The vector-returning to_cbor/to_msgpack/to_ubjson/to_bjdata/to_bson grew
the output buffer purely by geometric reallocation. Reserving an estimate
up front avoids the early reallocations, which is the dominant per-byte
cost for array/object-heavy output.
The estimate (binary_reserve_hint) is deliberately conservative and safe
against untrusted input: it consults only the top-level element count
(O(1), no walk of the DOM), guards the multiplication against overflow,
and clamps the result to a fixed 1 MiB ceiling, so a large or hostile DOM
can never force an oversized allocation here. The buffer still grows
geometrically past the hint, so an underestimate only costs a few later
reallocations; scalars/strings/binary are written in one shot and get no
hint. Reserving capacity does not change the bytes produced.
Throughput (g++/clang -O3, vs the previous commit):
cbor int array +10% / +13%
cbor object array +20% / +38%
Output is byte-for-byte identical to develop across the binary
differential corpus.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Address review findings on the binary writer output sinks
- binary_reserve_hint(): the 4-bytes-per-element estimate over-reserved by up
to 4x for arrays of small scalars (CBOR encodes 0..23 in one byte), and the
returned vector kept that capacity. Make the hint a strict lower bound on the
encoded size instead, which also removes the 1 MiB clamp whose branch no test
could reach (the largest container in the suite has 65793 elements).
- Guard the -Wduplicated-branches pragma with __GNUC__ >= 7. The warning does
not exist before GCC 7, so naming it made GCC 4.8/4.9/5/6 - which the CI
matrix still builds - warn under -Wpragmas on every including translation
unit, breaking downstream -Werror builds.
- Constrain the adapter constructor of binary_writer with the enable_if its
documentation already claimed, so a writer over some other sink type is no
longer advertised as constructible from an output adapter.
- Let output_vector_adapter wrap output_vector_sink rather than duplicating the
append logic, so the type-erased and templated paths share one implementation.
- Collapse the three copies of the memcpy/byte_swap/memcpy dance into a single
byte_swap_buffer() helper, and add the MSVC _byteswap_* intrinsics so MSVC no
longer falls back to the scalar shuffle this change exists to eliminate.
- Add a vector_writer() helper for the five vector-returning to_* overloads
instead of spelling out the writer type at each call site, and drop a dead
default member initializer on output_adapter_sink.
- New tests: the vector sink and the adapter sink must produce identical bytes
for every format (the two to_* overloads no longer delegate to each other and
could otherwise drift), and binary_reserve_hint() must never exceed the size
actually written.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Route the -Wduplicated-branches pragma through Hedley
Match #5485, which moved the binary writer's hand-rolled diagnostic
pragmas onto JSON_HEDLEY_PRAGMA (merged into develop while this branch
was open). The devirtualization's -Wduplicated-branches suppression in
write_compact_float was the one raw '#pragma GCC diagnostic' left; it
now uses JSON_HEDLEY_PRAGMA like the adjacent -Wfloat-equal line, still
guarded to GCC >= 7 and non-clang (the warning exists only there).
Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* Reject MessagePack/BSON binary subtypes that don't fit their wire format
Both formats store byte_container_with_subtype's subtype (a uint64_t)
in a single byte. The writers cast to std::int8_t/std::uint8_t without
a range check, so subtypes above 255 were silently truncated modulo
256 instead of raising an error. Throw out_of_range.413 instead when
the subtype exceeds the representable range of 0-255.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move the new binary-subtype regression test out of unit-regression2.cpp
unit-regression2.cpp is already at the edge of what the MinGW linker
can relocate; adding this test's ~26 lines tips test-regression2_cpp20
(clang, Windows) over into "relocation truncated to fit:
IMAGE_REL_AMD64_REL32 against `.rdata'" (see 8ce64b9c1 / b82717c8a for
the same failure mode). Split the test along format lines instead:
MessagePack assertions move to unit-msgpack.cpp, BSON assertions to
unit-bson.cpp. The CBOR round-trip guard is dropped as redundant --
unit-cbor.cpp's "Tagged values" section already round-trips subtypes
up to 8589934590, far past the 70000 checked here.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The array-range insert() overload checked that pos fits the current
value and that first/last share the same owning value, but never
verified that value is itself an array. Passing iterators from an
object, a primitive, or null handed value-initialized (singular)
std::vector iterators straight to array_t::insert(), which is
undefined behavior. Add the missing is_array() check, mirroring the
equivalent check already present in the object-range insert()
overload.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>