diff --git a/analysis/06-The_Myth_of_Clean_Input/readme.md b/analysis/06-The_Myth_of_Clean_Input/readme.md new file mode 100644 index 0000000..79be012 --- /dev/null +++ b/analysis/06-The_Myth_of_Clean_Input/readme.md @@ -0,0 +1,742 @@ +# Analysis #06 — The Myth of Clean Input + +## Problem + +Most algorithmic problems begin in roughly the same way: + +> Given an array. + +> Given a linked list. + +> Given a vector of temperatures. + +For example: + +```cpp +std::vector temperatures; +``` + +The task may then ask us to find the maximum value, calculate an average, remove duplicates, or perform some other operation. + +The input is assumed to: + +- already exist; +- have the expected type; +- use the expected representation; +- be free from corruption; +- contain values within valid ranges; +- be ready for use. + +This assumption feels so natural that it is rarely noticed. + +In real engineering, however, a clean object is rarely the starting point. + +More often, it is the final result of a long processing chain. + +--- + +## Typical Interview Thinking + +Consider a simple problem: + +> Find the maximum temperature. + +The candidate receives a ready-to-use container: + +```cpp +std::vector temperatures; +``` + +The solution may be almost trivial: + +```cpp +const auto max_temperature = + std::max_element(temperatures.begin(), temperatures.end()); +``` + +From there, the discussion may cover: + +- computational complexity; +- memory usage; +- empty input handling; +- use of the standard library; +- possible optimizations. + +All attention is focused on the algorithm. + +But one question is almost never asked: + +> Where did this `std::vector` come from? + +Who received the original data? + +Who verified the frame? + +Who determined the format? + +Who converted the raw sensor value into degrees? + +Who decided that the resulting number could be trusted? + +The interview starts with an already prepared object. + +A real system must first create that object. + +--- + +## Clean Input Is Not a Starting Condition + +Consider a temperature received from a remote sensor. + +At the business-logic level, it may look like this: + +```text +23.7 °C +``` + +But the system may have originally received nothing more than a sequence of bytes: + +```text +02 03 00 1A FF 7C 91 4D +``` + +Before those bytes can become a temperature, the data must pass through several stages: + +```text +UART / CAN / TCP + │ + ▼ +Receive raw bytes + │ + ▼ +Extract a complete frame + │ + ▼ +Validate frame length + │ + ▼ +Verify checksum + │ + ▼ +Check protocol version + │ + ▼ +Deserialize the payload + │ + ▼ +Identify the data source + │ + ▼ +Convert byte order + │ + ▼ +Apply scale and offset + │ + ▼ +Check the physical range + │ + ▼ +Normalized temperature + │ + ▼ +Append to std::vector + │ + ▼ +Find the maximum value +``` + +The maximum-value algorithm is the final step and may be the simplest step in the entire chain. + +--- + +## The Cost of Clean Input + +This declaration looks simple: + +```cpp +std::vector temperatures; +``` + +But that simplicity was not free. + +Before business logic can receive such a container, the system may already have had to: + +- receive data from an external source; +- identify message boundaries; +- detect corruption; +- check protocol-version compatibility; +- parse a binary representation; +- handle byte order; +- apply scaling; +- recognize reserved or unavailable values; +- verify physical plausibility; +- convert the result into an internal representation. + +The algorithm may require one line of code. + +The infrastructure that makes that line meaningful may require thousands. + +This is why business logic is often simpler than the code surrounding it. + +The formula may already be known. + +The algorithm may already exist in the standard library. + +The real work is ensuring that the transition from the external world to the object expected by that algorithm is correct. + +--- + +## Normalization: Valid Data Can Still Be Incomparable + +Not every input problem is caused by corruption. + +Sometimes each value is individually valid, but several values represent the same entity in different forms. + +Consider a task that searches for duplicate vehicle identifiers. + +An interview problem may provide this input: + +```text +ABC123 +ABC123 +ABC123 +``` + +The result is obvious. + +A real system may receive: + +```text +ABC123 +abc123 +ABC-123 + ABC123 +ABC123 +``` + +At the string level, these values are different. + +A duplicate-search algorithm will correctly report that they do not match. + +At the domain level, however, they may represent the same object. + +Before searching for duplicates, the system must define a canonical representation: + +- Is character case significant? +- Are separators meaningful? +- Should surrounding whitespace be removed? +- Which characters are permitted? +- Is there a canonical format? +- What should happen when the input is ambiguous? + +After normalization, the values may become: + +```text +ABC123 +ABC123 +ABC123 +ABC123 +ABC123 +``` + +Only now is the duplicate-search algorithm solving the correct problem. + +Before normalization, it was comparing representations rather than entities. + +--- + +## Normalization Is Part of the System Model + +Normalization can look like little more than string cleanup. + +In reality, it expresses domain rules. + +For example: + +- letter case may be irrelevant for one identifier and essential for another; +- two file paths may refer to the same object while remaining different strings; +- phone numbers may contain different country prefixes and formatting; +- MAC addresses may use different separators; +- timestamps may use different time zones; +- measurements may use different units; +- sensor values may require calibration. + +Normalization does not merely answer: + +> How should this string be modified? + +It answers: + +> What does this system consider to be the same value? + +There is no universal normalization procedure. + +It depends on the protocol, the contract, and the meaning of the data. + +--- + +## Not All Well-Formed Data Is Usable + +Successfully parsing a message does not mean that its contents are safe to use. + +Consider these temperatures: + +```text +23.7 +-40.0 +65535 +NaN +-273.15 +``` + +Every one of these values may be successfully represented as a number. + +Their meanings, however, are very different. + +`23.7` may be a normal measurement. + +`-40.0` may be valid, or it may be the lower limit of the sensor. + +`65535` may represent unavailable data. + +`NaN` may have appeared after an invalid calculation. + +`-273.15` is numerically valid, but for a particular device it almost certainly indicates a problem. + +This reveals several different levels of correctness. + +### Structural Correctness + +Can the message be parsed? + +### Protocol Correctness + +Does it conform to the expected protocol format and version? + +### Numeric Correctness + +Can the value be represented using the required type? + +### Semantic Correctness + +Does the value make sense within the domain? + +Syntactic validity does not guarantee meaningful data. + +--- + +## From Raw Data to a Trusted Object + +It is useful to view input handling not as one large validation step, but as a sequence of state transitions. + +```text +Raw bytes + │ + ▼ +Framed data + │ + ▼ +Integrity-checked frame + │ + ▼ +Parsed message + │ + ▼ +Normalized values + │ + ▼ +Semantically valid object + │ + ▼ +Trusted domain object + │ + ▼ +Business algorithm +``` + +At every stage, the system gains stronger guarantees. + +Raw bytes promise almost nothing. + +After framing, the message boundaries are known. + +After integrity checks, there is evidence that the data was not accidentally corrupted. + +After parsing, typed fields exist. + +After normalization, values use a consistent representation. + +After semantic checks, the object is known to be acceptable within the domain. + +Only then can the data be treated as trusted by a particular layer of the system. + +--- + +## Trust Must Be Local + +This leads to an important architectural principle: + +> Data is not simply trusted or untrusted. + +It is trusted only relative to a particular contract. + +A transport layer may guarantee that: + +- the complete frame was received; +- the checksum matches; +- the length is valid. + +It cannot guarantee that a temperature is physically meaningful. + +A parser may guarantee that: + +- message fields were extracted successfully; +- their sizes and types match the protocol. + +It cannot determine whether the value is acceptable for a specific device model. + +That responsibility belongs to another layer. + +Each layer checks its own invariants and passes a stronger representation to the next one. + +--- + +## Every Layer Earns Trust for the Next One + +A clean object at an algorithm boundary is not a magical property of the data. + +It is the result of fulfilled contracts. + +One layer says: + +> I verified the integrity of the frame. + +The next says: + +> I parsed the message according to a supported protocol version. + +The next says: + +> I converted the values into the system's internal units. + +The next says: + +> I confirmed that the object is valid within this domain. + +Only then may the business logic assume: + +> This is a valid temperature. + +That assumption is not justified because the external world is reliable. + +It is justified because the previous layers did their work. + +--- + +## Why Not Validate Everything Everywhere? + +Distrusting input can lead to another bad conclusion: + +> Every function should repeat every validation step. + +That creates different problems: + +- duplicated logic; +- contradictory checks; +- unclear ownership of responsibilities; +- more complex code; +- uncertainty about which guarantees already exist. + +A function that accepts raw bytes must not assume that they are safe. + +A function that accepts an object which can only be created after successful verification does not need to repeat the entire process. + +Good architecture does not eliminate trust. + +It makes the origin of trust explicit. + +--- + +## Types as Evidence of the Path Already Taken + +One practical way to express this is to use different types for different processing stages. + +Instead of passing the same generic object through the entire system, the stages can be represented explicitly: + +```cpp +struct RawFrame; +struct VerifiedFrame; +struct ParsedTemperatureMessage; +struct NormalizedTemperature; +``` + +The interfaces can then reflect the available guarantees: + +```cpp +std::optional +verify_frame(const RawFrame& frame); + +std::optional +parse_message(const VerifiedFrame& frame); + +std::optional +normalize_temperature(const ParsedTemperatureMessage& message); +``` + +Business logic can accept only the normalized value: + +```cpp +void process_temperature(const NormalizedTemperature& temperature); +``` + +This does not make the data absolutely true. + +It makes the stages already completed explicit. + +It also prevents raw input from being passed accidentally into code that expects a verified object. + +--- + +## What Happens When Processing Fails? + +Data evolution does not always end with a valid business object. + +Every stage may reject the input: + +```text +Raw bytes + │ + ├── incomplete frame + ├── unsupported version + ├── invalid checksum + ├── malformed payload + ├── unknown sensor + ├── invalid scaling + ├── out-of-range value + └── valid temperature +``` + +This introduces another major part of real engineering that is usually absent from algorithmic problems: + +- the message may need to be discarded; +- the failure may need to be logged; +- a diagnostic counter may need to be incremented; +- the source may need to be reconnected; +- the system may need to use the last known valid value; +- a component may enter a degraded mode; +- the failure may affect safety-related behavior. + +In an interview problem, an invalid value is often just an edge case. + +In a real system, it may trigger an entirely different operating scenario. + +--- + +## The Algorithm Still Matters + +None of this means that algorithms are unimportant. + +Once data has been converted into a correct internal model, the algorithm still needs to be: + +- correct; +- efficient; +- understandable; +- appropriate for the system constraints. + +The problem begins when solving a task over a clean array is treated as a complete model of engineering ability. + +An algorithm solves a problem under a set of assumptions. + +An engineer must also: + +- discover those assumptions; +- determine whether they are valid; +- assign responsibility for enforcing them; +- express the resulting guarantees in interfaces and architecture. + +--- + +## What This Actually Tests + +A problem over a ready-made container can test: + +- knowledge of data structures; +- algorithmic reasoning; +- complexity analysis; +- recognition of known patterns; +- implementation accuracy. + +It says much less about a candidate's ability to: + +- work with external data sources; +- design trust boundaries; +- parse protocols; +- normalize representations; +- define semantic validity; +- design diagnostics; +- handle partial failures; +- create reliable contracts between layers. + +This does not make the algorithmic task useless. + +It only limits what can reasonably be concluded from it. + +--- + +## Where the Interview Ends and Engineering Begins + +An interview problem often presents this model: + +```text +Clean input + │ + ▼ +Algorithm + │ + ▼ +Result +``` + +A real system often looks more like this: + +```text +Physical world + │ + ▼ +Electrical signal + │ + ▼ +Raw bytes + │ + ▼ +Transport framing + │ + ▼ +Integrity checks + │ + ▼ +Protocol parsing + │ + ▼ +Version handling + │ + ▼ +Normalization + │ + ▼ +Semantic validation + │ + ▼ +Domain object + │ + ▼ +Algorithm + │ + ▼ +System decision +``` + +The interview begins near the end of this chain. + +Engineering is responsible for the entire chain. + +--- + +## The Evolution of Data + +We can now return to the temperature example. + +Initially, the system does not have a temperature. + +It has a signal. + +Then it has bytes. + +Then a frame. + +Then a message. + +Then a raw sensor value. + +Then a value expressed in physical units. + +Then a normalized and semantically valid measurement. + +Only after all of that does a number appear that can safely be stored in a container and passed to an algorithm. + +```text +Signal + ↓ +Bytes + ↓ +Frame + ↓ +Verified frame + ↓ +Parsed message + ↓ +Raw sensor value + ↓ +Calibrated value + ↓ +Normalized temperature + ↓ +Trusted domain object + ↓ +std::vector + ↓ +std::max_element +``` + +The maximum-search algorithm does not create the meaning of the data. + +It consumes meaning that was established by the previous layers. + +--- + +## Key Takeaway + +Clean input is not a starting point. + +It is an engineering result. + +It exists only after the system has: + +- identified the structure of the data; +- verified its integrity; +- understood its format; +- converted it into a canonical representation; +- checked its meaning; +- established a contract of trust. + +The engineer's first question is therefore not: + +> How do I process this array? + +It is: + +> Why can this array be trusted? + +And then: + +> Which layer guarantees that? + +--- + +## Project Perspective + +> Exists in real engineering? +> Yes. Almost constantly. + +> Exists in interview form? +> Usually not. Most of the data journey is hidden by the problem statement. + +Algorithmic tasks are useful for evaluating work on already prepared structures. + +But they usually begin with a result that a real system still has to produce. + +That is the myth of clean input: + +> Data does not arrive ready for the algorithm. + +> Engineering makes it ready.