A C++ program that accepts 42extra as 42 has parsed a prefix, not validated a number. That difference matters when the value controls a file offset, an allocation, or a permission decision. At each input boundary, define which bytes are allowed, how they become a typed value, and what to do if any input remains unparsed.
Input might come from a command line, configuration file, network message, or local data file. None guarantees a value that matches the program's expectations. Validation gives internal code a value it can rely on. It does not replace memory-safe design or authorization checks, but it keeps malformed and out-of-range values from silently reaching sensitive operations.
Write the input contract before the parser
Suppose a tool accepts an ASCII decimal count from 1 through 1000. “Integer required” is not enough of a contract. Specify that the input contains only bytes 0 through 9; reject signs, spaces, separators, and trailing text; require the parser to consume the entire input; and reject zero or anything above 1000. Limit the raw input size too. A million-digit string is invalid even if the numeric parser would eventually reject it.
Each check belongs at a particular stage:
- Framing: Is the complete record available, and is its byte length within the limit?
- Syntax: Does the record use the permitted characters and grammar?
- Conversion: Can the representation fit in the chosen C++ type?
- Domain: Is the converted value allowed for this operation?
- Context: Is the action permitted in the current state and for this caller?
A count of 900 may pass syntax and range checks but still be too large when only 20 entries remain. A well-formed filename does not establish permission to open the file, either. Put contextual checks near the operation they protect, where the relevant state is available.

Bound input before parsing it
A parser cannot enforce a useful resource limit if the program has already read unlimited data into memory. Set maximum record and field sizes at the entry point. For length-prefixed messages, check the declared length against a fixed limit before allocating or reading the body. For line-based input, use a bounded read strategy: std::getline can keep growing its destination string while it searches for a newline. If a line is too long, consume or close the rest of that record according to the protocol. Otherwise, the next read may treat the remainder as a new message.
Byte limits and character limits are different. UTF-8 characters can occupy multiple bytes, so a text field's visible length may differ from its byte length. Say which measure the interface uses. At a network boundary, a byte limit is usually the first defense against excessive memory use; apply any character-count rule after valid text decoding.
Parse a complete decimal value
For a narrow ASCII integer field, std::from_chars converts without locale-dependent whitespace handling or exceptions. This function applies the contract above, including a small raw-length cap. It returns no value on failure so the caller can report an appropriate error.
#include <charconv>
#include <cstddef>
#include <optional>
#include <string_view>
#include <system_error>
std::optional<unsigned> parse_count(std::string_view text) {
if (text.empty() || text.size() > 4) {
return std::nullopt;
}
for (char ch : text) {
if (ch < '0' || ch > '9') {
return std::nullopt;
}
}
unsigned value{};
const char* first = text.data();
const char* last = first + text.size();
auto [end, error] = std::from_chars(first, last, value, 10);
if (error != std::errc{} || end != last) {
return std::nullopt;
}
if (value < 1 || value > 1000) {
return std::nullopt;
}
return value;
}
The four-byte cap admits every valid value, including 1000, but it does not replace the range check: 9999 is also four bytes. The digit loop states the grammar independently of conversion-library behavior and rejects embedded NUL bytes as well as punctuation. Checking end == last guards against accepting a partial parse if the grammar changes later.
For a signed field, decide whether to allow a leading minus sign and whether -0 is acceptable. Do not assume parsing into an unsigned type expresses a “nonnegative integer” policy consistently across conversion APIs. You can use std::stoi, but account for its whitespace acceptance, exceptions, and position output. Avoid atoi for security-sensitive validation: it cannot reliably distinguish invalid input from a valid zero.
Validate types and arithmetic separately
A valid number can still lead to an unsafe calculation. Overflow or narrowing may happen after parsing, especially when a value determines a buffer size. Check a requested count against the container's limits before multiplying or converting it:
#include <cstddef>
#include <limits>
#include <vector>
bool fits_byte_buffer(std::size_t count, std::size_t width,
const std::vector<char>& buffer) {
if (width == 0 || count > buffer.max_size() / width) {
return false;
}
return count * width <= buffer.max_size();
}
Dividing before multiplying avoids calculating an overflowing product. In practice, an application's memory limit should usually be far below max_size(), which reports a container limit rather than an acceptable workload. Reject negative signed lengths before casting them to std::size_t. Before passing a std::size_t length to an API with a narrower integer type, compare it with that type's maximum. Parsing a value and using it safely are separate jobs.
Define text rules precisely
“Alphanumeric” needs a character-set definition. An ASCII identifier in an internal configuration file might allow letters, digits, underscores, and hyphens, up to 32 bytes. Explicit ASCII byte comparisons keep that rule predictable across locales. If you use std::isalnum or another function from <cctype>, convert the value to unsigned char first (or pass EOF where applicable). A negative signed char causes undefined behavior, and locale-dependent classification may not match a strict ASCII contract.
ASCII-only rules are often wrong for international names or free-form text. Decode using the encoding the interface promises, such as UTF-8, and reject malformed text. Then apply field-specific rules for control characters, length, and any normalization needed for comparison. Unicode normalization and case folding are not byte-by-byte lowercase conversion. If two spellings must identify the same account, define that rule centrally and test it so components do not disagree.
Do not silently rewrite input unless the contract says you will. Trimming spaces, dropping invalid bytes, or replacing characters can make distinct inputs collide. Rejecting unexpected whitespace is often clearer for a machine-readable protocol. Trimming a user-entered display name may be reasonable if the behavior is specified and later checks use the resulting value.

Keep structure, meaning, and output safety apart
A filename shows why one generic “sanitize” function rarely works. An upload label might use a restricted character set and forbid path separators. A program that opens user-selected paths needs another approach: choose an approved base directory, handle paths according to the platform, and enforce access controls. Stripping .. substrings does not establish where a resolved path points. Filesystem links and platform-specific path forms complicate that check, and validating a path string does not remove races between checking and opening it.
The same distinction applies to database queries and displayed text. Pass a validated username to the database through a parameterized query, not SQL string concatenation. Encode text for its destination HTML context even if it passed input validation. A field allowlist cannot protect against every interpreter that might later receive the value.
Use structured parsing for structured input
If an application accepts JSON, CSV, or another defined format, use a parser that understands its grammar instead of splitting on punctuation. CSV fields can contain quoted delimiters; JSON strings can contain escaped characters. Once parsing produces a structure, validate the fields: required members, types, allowed values, nesting depth, collection sizes, and the policy for duplicate fields where the format and parser permit them.
Treat missing, empty, and null values separately when they mean different things. A missing optional setting may take a documented default; an explicitly empty value may be invalid. Never turn a failed parse into a successful privileged setting through a default. Reject unknown fields if they could hide typos or cause different versions to interpret a record differently. If forward compatibility calls for accepting them, document and test that choice.
Return errors without losing control of the boundary
A validation function should produce a value that satisfies its contract or report failure. std::optional is enough when callers need only success or failure. An interactive tool may benefit from bounded error categories such as “too long,” “invalid character,” and “outside allowed range.” Give useful external errors without echoing sensitive input into logs. When logging is needed, record the field name, failure category, and a safe request identifier—not secrets or arbitrarily long attacker-controlled strings.
Keep partially validated state out of the rest of the application. Parse a whole record into temporary values, check cross-field constraints, and then commit one valid object. For a buffer slice, validate both offset and length: check offset <= buffer.size(), then length <= buffer.size() - offset. That avoids overflow in offset + length. If the buffer can change concurrently, protect the check and the use with appropriate synchronization; a check against an old size cannot protect a later access.
Test the rejected cases as carefully as the accepted ones
Derive tests from the contract, not just typical examples. For parse_count, accept 1 and 1000. Reject an empty string, 0, 1001, 9999, a leading space, either sign, a decimal point, trailing letters, an embedded NUL byte, and a five-byte digit string. Build embedded-NUL cases with an explicit length, such as std::string_view(data, length). A C-string constructor would stop at the first NUL and test something else.
| Boundary | Useful test | Expected behavior |
|---|---|---|
| Length | One byte above the maximum | Reject before costly parsing |
| Syntax | Valid prefix plus extra data | Reject the entire field |
| Range | Minimum and maximum, then one beyond each | Accept only values inside the contract |
| Arithmetic | Count near the allocation limit | Reject before multiplication or narrowing |
| Encoding | Malformed byte sequence in a text field | Reject or handle under a documented policy |
Fuzz testing can supplement named cases by feeding generated byte sequences to a parser in a local test environment. Check more than whether it crashes: every successful result must satisfy the field's length, syntax, and range rules. Run relevant compiler sanitizers during tests to catch memory and arithmetic errors around the parser. A sanitizer cannot decide whether 42extra is valid; the contract and its tests must do that.
When you add a field, put its grammar and size limit beside the code that reads it. For a 32-byte ASCII identifier, submit 32 permitted bytes and then 33 in a regression test. Verify that the first value passes and the second is rejected before it reaches storage.
