mirror of
https://github.com/nlohmann/json.git
synced 2026-10-08 00:35:14 +07:00
* Check in the fuzzers that parsing without exceptions agrees Each fuzzer driver now also parses its input with allow_exceptions = false. That call must never throw a parse_error, must return a discarded value where parsing with exceptions fails, and must return the same value where it succeeds. Values are compared by their dump(), because NaN is not equal to itself. A plain !is_discarded() assertion, as suggested in #3642, would never fail: the drivers parse with exceptions, so a result can never be discarded. tests/fuzzing.md describes the checks and notes that OSS-Fuzz and CIFuzz already run LeakSanitizer, because their default address sanitizer includes it. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Use JSON_HAS_RANGE_VIEW_CONVERSION in the range view regression tests #5728 combined the JSON_HAS_RANGES and MinGW conditions into JSON_HAS_RANGE_VIEW_CONVERSION, but three test guards still spelled them out. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Test the serializer's buffers at their boundaries The dump() indent overflow survived full line coverage because the tests grew its buffer by only one step. This adds tests that land exactly on, and one past, the limits of the other two serializer buffers: - write_buffer (1024 bytes): strings of 1023, 1024 and 1025 bytes at the top level, and of 1022 and 1023 bytes inside an array, so that both guards in put_string() are hit at their boundary. Each is checked for dump() and for stream output. - string_buffer (512 bytes, flushed when fewer than 13 bytes remain): runs of two-byte escapes, and a surrogate pair written with 14 bytes of room, right after a flush, and one escape later. - The 8-byte bulk scan from the serializer side: 0 to 17 plain bytes followed by a quote, a control character, or a non-ASCII character. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Test the chunked string and binary reads of all binary formats The binary readers read strings and binary values in chunks of 4096 bytes. Only CBOR tested lengths around that size. MessagePack, UBJSON, BJData and BSON now round-trip lengths 0, 1, 4095, 4096, 4097, 8192 and 100000 from vector and pointer input, and must report a truncated payload as a parse error. UBJSON reads binary values as arrays of numbers, so it is tested with strings only. BJData binary values reach the chunked read only in Draft 3. BON8 decodes strings byte by byte and does not use this path. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
131 lines
4.6 KiB
C++
131 lines
4.6 KiB
C++
// __ _____ _____ _____
|
|
// __| | __| | | | JSON for Modern C++ (supporting code)
|
|
// | | |__ | | | | | | version 3.12.0
|
|
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
|
|
//
|
|
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
|
|
// SPDX-License-Identifier: MIT
|
|
|
|
/*
|
|
This file implements a parser test suitable for fuzz testing. Given a byte
|
|
array data, it performs the following steps:
|
|
|
|
- j0 = from_ubjson(data, allow_exceptions = false)
|
|
- j1 = from_ubjson(data)
|
|
- assert(j0 is discarded if parsing j1 fails, and j0 == j1 otherwise)
|
|
- vec2 = to_ubjson(j1, use_size = false, use_type = false)
|
|
- vec3 = to_ubjson(j1, use_size = true, use_type = false)
|
|
- vec4 = to_ubjson(j1, use_size = true, use_type = true)
|
|
- j2 = from_ubjson(vec2)
|
|
- j3 = from_ubjson(vec3)
|
|
- j4 = from_ubjson(vec4)
|
|
- assert(to_ubjson(j2, use_size = false, use_type = false) == vec2)
|
|
- assert(to_ubjson(j3, use_size = true, use_type = false) == vec3)
|
|
- assert(to_ubjson(j4, use_size = true, use_type = true) == vec4)
|
|
|
|
The unit tests run the same checks on a fixed corpus (see the "UBJSON round-trip
|
|
invariants" test case), so keep both in sync.
|
|
|
|
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
|
|
drivers.
|
|
*/
|
|
|
|
#include <cassert>
|
|
#include <nlohmann/json.hpp>
|
|
|
|
// the round-trip checks below are assertions; NDEBUG would compile them away
|
|
#ifdef NDEBUG
|
|
#error "the fuzzer drivers must be built without NDEBUG"
|
|
#endif
|
|
|
|
using json = nlohmann::json;
|
|
|
|
// compares dumps rather than values, because NaN != NaN; keep writes strings
|
|
// byte for byte, so ill-formed UTF-8 that a binary reader accepts cannot throw
|
|
static bool same_value(const json& lhs, const json& rhs)
|
|
{
|
|
return lhs.dump(-1, ' ', false, json::error_handler_t::keep) == rhs.dump(-1, ' ', false, json::error_handler_t::keep);
|
|
}
|
|
|
|
// see http://llvm.org/docs/LibFuzzer.html
|
|
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
|
|
{
|
|
std::vector<uint8_t> const vec1(data, data + size);
|
|
|
|
// step 0: parse input without exceptions; a parse error must then be
|
|
// reported as a discarded value, never thrown
|
|
json j_noexcept;
|
|
bool noexcept_threw = false;
|
|
try
|
|
{
|
|
j_noexcept = json::from_ubjson(vec1, true, false);
|
|
}
|
|
catch (const json::parse_error&)
|
|
{
|
|
assert(false);
|
|
}
|
|
catch (const json::exception&)
|
|
{
|
|
// type and out-of-range errors are not parse errors and still throw
|
|
noexcept_threw = true;
|
|
}
|
|
// whether step 1 succeeded; if not, the catch blocks below check that
|
|
// step 0 failed, too
|
|
bool parsed = false;
|
|
|
|
try
|
|
{
|
|
// step 1: parse input
|
|
json const j1 = json::from_ubjson(vec1);
|
|
parsed = true;
|
|
|
|
// without exceptions, the same input must give the same value
|
|
assert(!noexcept_threw && !j_noexcept.is_discarded() && same_value(j_noexcept, j1));
|
|
|
|
try
|
|
{
|
|
// step 2.1: round trip without adding size annotations to container types
|
|
std::vector<uint8_t> const vec2 = json::to_ubjson(j1, false, false);
|
|
|
|
// step 2.2: round trip with adding size annotations but without adding type annotations to container types
|
|
std::vector<uint8_t> const vec3 = json::to_ubjson(j1, true, false);
|
|
|
|
// step 2.3: round trip with adding size as well as type annotations to container types
|
|
std::vector<uint8_t> const vec4 = json::to_ubjson(j1, true, true);
|
|
|
|
// parse serialization
|
|
json const j2 = json::from_ubjson(vec2);
|
|
json const j3 = json::from_ubjson(vec3);
|
|
json const j4 = json::from_ubjson(vec4);
|
|
|
|
// serializations must match
|
|
assert(json::to_ubjson(j2, false, false) == vec2);
|
|
assert(json::to_ubjson(j3, true, false) == vec3);
|
|
assert(json::to_ubjson(j4, true, true) == vec4);
|
|
}
|
|
catch (const json::parse_error&)
|
|
{
|
|
// parsing a UBJSON serialization must not fail
|
|
assert(false);
|
|
}
|
|
}
|
|
catch (const json::parse_error&)
|
|
{
|
|
// parse errors are ok, because input may be random bytes
|
|
assert(parsed || noexcept_threw || j_noexcept.is_discarded());
|
|
}
|
|
catch (const json::type_error&)
|
|
{
|
|
// type errors can occur during parsing, too
|
|
assert(parsed || noexcept_threw || j_noexcept.is_discarded());
|
|
}
|
|
catch (const json::out_of_range&)
|
|
{
|
|
// out of range errors may happen if provided sizes are excessive
|
|
assert(parsed || noexcept_threw || j_noexcept.is_discarded());
|
|
}
|
|
|
|
// return 0 - non-zero return values are reserved for future use
|
|
return 0;
|
|
}
|