Compare commits
15 Commits
90793c9931
...
02-Update_
| Author | SHA1 | Date | |
|---|---|---|---|
| 2da217c0ce | |||
| 37583d820a | |||
| be12197ac0 | |||
| 155cca5eeb | |||
| 7a74d9fca3 | |||
| 2b49cbe010 | |||
| 556163a92b | |||
| cf2ea7f478 | |||
| 5b8566e69c | |||
| 72175f4fdb | |||
| 7a5e040c38 | |||
| 415c9757a6 | |||
| d95dd32f36 | |||
| d047f24112 | |||
| 37c92efd80 |
1
.gitignore
vendored
Normal file
1
.gitignore
vendored
Normal file
@@ -0,0 +1 @@
|
||||
.tmp/
|
||||
87
adr/ADR-001-repository-name.md
Normal file
87
adr/ADR-001-repository-name.md
Normal file
@@ -0,0 +1,87 @@
|
||||
# ADR-001: Use `ota-reference-design` as the Repository Name
|
||||
|
||||
* **Status:** Accepted
|
||||
* **Date:** 2026-07-29
|
||||
|
||||
## Context
|
||||
|
||||
The project requires a repository name that communicates its purpose clearly without tying the design to a single hardware platform, implementation language, or update framework.
|
||||
|
||||
The repository will contain:
|
||||
|
||||
* architectural documentation;
|
||||
* OTA design decisions;
|
||||
* diagrams;
|
||||
* reference implementations;
|
||||
* update tooling;
|
||||
* failure-injection experiments;
|
||||
* security and reliability analysis.
|
||||
|
||||
The initial practical implementation may use Raspberry Pi hardware, but the overall design should remain applicable to other embedded Linux platforms.
|
||||
|
||||
The repository is not intended to provide a production-ready OTA framework or a reusable software library. Its primary purpose is to demonstrate and explain the architecture of a reliable and secure OTA update system.
|
||||
|
||||
## Decision
|
||||
|
||||
The repository will be named:
|
||||
|
||||
```text
|
||||
ota-reference-design
|
||||
```
|
||||
|
||||
The repository will initially remain a standalone top-level repository rather than being placed under a broader group such as `engineering`, `embedded-systems`, or `systems-engineering`.
|
||||
|
||||
## Alternatives Considered
|
||||
|
||||
### `embedded-ota-lab`
|
||||
|
||||
This name emphasizes experimentation and practical work.
|
||||
|
||||
It was not selected because the project will contain more than experiments. It will also document architecture, security, reliability, design decisions, and implementation trade-offs.
|
||||
|
||||
### `ota-update-lab`
|
||||
|
||||
This name is simple but somewhat redundant because OTA already means over-the-air updating.
|
||||
|
||||
It also makes the project sound more like a temporary collection of experiments than a structured reference design.
|
||||
|
||||
### `reliable-ota`
|
||||
|
||||
This name emphasizes one of the primary goals of the project.
|
||||
|
||||
It was not selected because it does not clearly communicate that the repository is educational and architectural rather than a production-ready OTA product.
|
||||
|
||||
### A Repository Under an `engineering` Group
|
||||
|
||||
A broader parent group could eventually contain several related repositories.
|
||||
|
||||
It was not selected at this stage because a group containing only one repository adds unnecessary hierarchy and does not yet provide meaningful organization.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
* The name clearly communicates that the project is an OTA reference design.
|
||||
* The repository is not tied to Raspberry Pi or any other specific platform.
|
||||
* The name allows the project to include documentation, code, diagrams, and experiments.
|
||||
* Readers are less likely to mistake the project for a production-ready framework.
|
||||
* The name remains appropriate if additional hardware platforms are added later.
|
||||
|
||||
### Negative
|
||||
|
||||
* The name is broader than the initial Raspberry Pi implementation.
|
||||
* Readers may still require the README to understand the exact project scope.
|
||||
* The term “reference design” may suggest a more complete implementation than exists during the early stages of development.
|
||||
|
||||
## Notes
|
||||
|
||||
A broader repository group may be introduced later if several related systems-engineering projects are created.
|
||||
|
||||
Possible future group names include:
|
||||
|
||||
```text
|
||||
embedded-systems
|
||||
systems-engineering
|
||||
```
|
||||
|
||||
Moving the repository into such a group would not require changing the repository name.
|
||||
87
adr/ADR-002-license.md
Normal file
87
adr/ADR-002-license.md
Normal file
@@ -0,0 +1,87 @@
|
||||
# ADR-002: Use the MIT License
|
||||
|
||||
* **Status:** Accepted
|
||||
* **Date:** 2026-07-29
|
||||
|
||||
## Context
|
||||
|
||||
The project is intended to be an open educational and engineering reference for designing OTA update systems on embedded Linux.
|
||||
|
||||
The repository may contain:
|
||||
|
||||
* documentation;
|
||||
* diagrams;
|
||||
* source code;
|
||||
* scripts;
|
||||
* configuration examples;
|
||||
* reference implementations;
|
||||
* test utilities;
|
||||
* failure-injection experiments.
|
||||
|
||||
The license should allow engineers to study, modify, reuse, and adapt the material in both personal and commercial projects.
|
||||
|
||||
The license should also be familiar, permissive, and easy to understand. The project does not require derivative works to remain open source.
|
||||
|
||||
## Decision
|
||||
|
||||
The repository will be licensed under the MIT License.
|
||||
|
||||
The license will apply to the repository content unless a specific file or third-party component explicitly states otherwise.
|
||||
|
||||
A standard `LICENSE` file containing the MIT License text will be included at the repository root.
|
||||
|
||||
## Alternatives Considered
|
||||
|
||||
### Apache License 2.0
|
||||
|
||||
Apache 2.0 is a permissive license that includes an explicit patent grant and patent retaliation provisions.
|
||||
|
||||
It was not selected because the project is currently an educational reference design rather than a large infrastructure framework or commercially governed software platform.
|
||||
|
||||
The additional legal complexity does not currently provide enough practical benefit over MIT.
|
||||
|
||||
### GNU General Public License
|
||||
|
||||
The GPL would require derivative works distributed under certain conditions to remain under the same license.
|
||||
|
||||
It was not selected because the project is intended to encourage broad reuse, including adoption of ideas and code in commercial embedded products.
|
||||
|
||||
A copyleft requirement could discourage some companies or engineers from using the examples.
|
||||
|
||||
### BSD 2-Clause or BSD 3-Clause License
|
||||
|
||||
The BSD licenses are permissive and would also be suitable for the project.
|
||||
|
||||
They were not selected because MIT is more familiar to many readers, concise, and already consistent with other projects maintained by the author.
|
||||
|
||||
### No Explicit License
|
||||
|
||||
Without a license, the repository would remain protected by copyright by default and other people would not have clear legal permission to reuse the material.
|
||||
|
||||
This would conflict with the educational and open-source goals of the project.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
* The project can be used in open-source and commercial environments.
|
||||
* Engineers can copy, modify, and redistribute examples.
|
||||
* The license is short and widely understood.
|
||||
* The project remains easy to adopt.
|
||||
* The license is consistent with the author’s other open-source repositories.
|
||||
* Attribution and preservation of the license notice are still required.
|
||||
|
||||
### Negative
|
||||
|
||||
* Modified versions are not required to remain open source.
|
||||
* Improvements may be used commercially without being contributed back.
|
||||
* The license does not include the explicit patent language provided by Apache 2.0.
|
||||
* The software is provided without warranty or liability protection beyond the license terms.
|
||||
|
||||
## Notes
|
||||
|
||||
The repository is a reference design and educational project.
|
||||
|
||||
The MIT License permits reuse but does not imply that the implementation is production-ready, certified, secure for every deployment, or suitable for safety-critical systems.
|
||||
|
||||
These limitations should also be explained in the repository README.
|
||||
120
adr/ADR-003-reference-design.md
Normal file
120
adr/ADR-003-reference-design.md
Normal file
@@ -0,0 +1,120 @@
|
||||
# ADR-003: Build a Reference Design, Not a Production Framework
|
||||
|
||||
* **Status:** Accepted
|
||||
* **Date:** 2026-07-29
|
||||
|
||||
## Context
|
||||
|
||||
OTA update systems differ significantly between products.
|
||||
|
||||
Their architecture depends on factors such as:
|
||||
|
||||
* hardware platform;
|
||||
* bootloader;
|
||||
* storage layout;
|
||||
* operating system;
|
||||
* security model;
|
||||
* network availability;
|
||||
* fleet size;
|
||||
* update frequency;
|
||||
* bandwidth limitations;
|
||||
* safety and regulatory requirements;
|
||||
* manufacturing and key-provisioning processes.
|
||||
|
||||
A generic production-ready OTA framework would need to support many combinations of hardware, bootloaders, update artifacts, deployment systems, security policies, and fleet-management requirements.
|
||||
|
||||
Building such a framework would significantly increase the project scope and could distract from its primary goal: explaining how reliable and secure OTA systems are designed.
|
||||
|
||||
The project should demonstrate engineering principles and concrete implementation choices without claiming universal applicability.
|
||||
|
||||
## Decision
|
||||
|
||||
The project will be developed as a practical OTA reference design.
|
||||
|
||||
It will provide:
|
||||
|
||||
* documented requirements;
|
||||
* architecture descriptions;
|
||||
* Architecture Decision Records;
|
||||
* update-state models;
|
||||
* security analysis;
|
||||
* reliability analysis;
|
||||
* diagrams;
|
||||
* example implementations;
|
||||
* failure-injection tests;
|
||||
* at least one working embedded Linux implementation.
|
||||
|
||||
The initial implementation may use Raspberry Pi as a practical demonstration platform, but the architecture will not be presented as Raspberry Pi-specific.
|
||||
|
||||
The project will not claim to be:
|
||||
|
||||
* a production-ready OTA framework;
|
||||
* a complete fleet-management service;
|
||||
* a universal update client;
|
||||
* a certified safety-critical solution;
|
||||
* a replacement for established commercial OTA platforms.
|
||||
|
||||
## Alternatives Considered
|
||||
|
||||
### Build a Reusable OTA Framework
|
||||
|
||||
A reusable framework could provide common APIs, backend services, bootloader integrations, and platform adapters.
|
||||
|
||||
It was not selected because it would require a much larger scope, long-term compatibility guarantees, extensive platform testing, and a stable public API.
|
||||
|
||||
It would also shift the project away from architecture and engineering analysis toward product development and maintenance.
|
||||
|
||||
### Build a Raspberry Pi-Specific OTA Tutorial
|
||||
|
||||
A platform-specific tutorial would be easier to implement and explain.
|
||||
|
||||
It was not selected because the main engineering principles—atomicity, rollback, artifact verification, health checks, boot control, and failure recovery—apply to many embedded Linux devices.
|
||||
|
||||
Tying the project too closely to Raspberry Pi would unnecessarily limit its usefulness.
|
||||
|
||||
### Publish Documentation Without a Working Implementation
|
||||
|
||||
A documentation-only project would allow broader architectural discussion without platform-specific complexity.
|
||||
|
||||
It was not selected because a working implementation is necessary to validate assumptions and demonstrate real failure modes.
|
||||
|
||||
Without practical experiments, the project could remain too theoretical.
|
||||
|
||||
### Build Only an OTA Update Client
|
||||
|
||||
A standalone client would provide a concrete software artifact.
|
||||
|
||||
It was not selected because OTA reliability depends on the complete system, including the bootloader, storage layout, update metadata, health checking, rollback, signing, and deployment process.
|
||||
|
||||
The client alone would not demonstrate the full architecture.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
* The project can focus on engineering reasoning and architecture.
|
||||
* Design decisions and trade-offs can be documented clearly.
|
||||
* The implementation can remain understandable and suitable for learning.
|
||||
* Platform-specific details can be isolated from general principles.
|
||||
* Additional reference implementations can be added later.
|
||||
* The project can evolve without promising a stable production API.
|
||||
* Failure scenarios can be explored openly without presenting the design as universally safe.
|
||||
|
||||
### Negative
|
||||
|
||||
* Users cannot assume the code is ready for direct production deployment.
|
||||
* Some components may be simplified for clarity.
|
||||
* Platform integration may require significant additional work.
|
||||
* The project may not cover manufacturing, provisioning, compliance, or very large fleet operations.
|
||||
* Readers must evaluate whether each design decision applies to their own product.
|
||||
* The boundary between a complete reference implementation and a framework must remain clearly documented.
|
||||
|
||||
## Notes
|
||||
|
||||
The repository README should contain a visible disclaimer similar to:
|
||||
|
||||
> This repository contains an educational OTA reference design for embedded Linux. It demonstrates one possible approach to reliable and secure system updates. It is not a production-ready framework and must be adapted, reviewed, tested, and hardened for each target product.
|
||||
|
||||
The project should prefer explicit architectural decisions over hidden assumptions.
|
||||
|
||||
When a design choice is platform-specific, it should be documented separately from the general OTA architecture.
|
||||
129
adr/ADR-004-full-system-image-updates.md
Normal file
129
adr/ADR-004-full-system-image-updates.md
Normal file
@@ -0,0 +1,129 @@
|
||||
# ADR-004: Full-System Image Updates
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
An OTA architecture must define the unit of update.
|
||||
|
||||
One possible approach is to treat the complete software stack as a single release artifact and distribute it as a prebuilt, tested system image.
|
||||
|
||||
For firmware-based devices, a full image is often the only practical update unit. For embedded Linux and other more complex platforms, it is one of several available strategies.
|
||||
|
||||
## Decision
|
||||
|
||||
Use full-system image updates when the device software is released and maintained as a single product, and when reproducibility, a known system state, and reliable system-level rollback are more important than minimizing update size or supporting independent component lifecycles.
|
||||
|
||||
## Decision Drivers
|
||||
|
||||
Full-system image updates are preferred when:
|
||||
|
||||
- the software stack is released as a single product;
|
||||
- the device must remain in one of a small number of predefined system states;
|
||||
- system-level rollback is required;
|
||||
- reproducibility is more important than minimizing update size;
|
||||
- the device fleet is sufficiently homogeneous;
|
||||
- image size and storage requirements are acceptable;
|
||||
- centralized release and validation of the complete system are practical.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- The image version identifies the version of the complete software system.
|
||||
- The device runs a predefined system state rather than an arbitrary combination of component versions.
|
||||
- The exact image installed on the device is the image that was validated before release.
|
||||
- Dependency resolution is not required during installation.
|
||||
- Diagnosis of the installed system version is simplified.
|
||||
- The approach integrates naturally with A/B layouts and system-level rollback.
|
||||
- The number of possible software configurations is reduced.
|
||||
- The software can be built, tested, signed, released, and rolled back as a single product.
|
||||
|
||||
### Negative
|
||||
|
||||
- Even a small change requires a new system image release.
|
||||
- Every image release requires full regression testing.
|
||||
- Unchanged components are transferred and written again.
|
||||
- Network traffic and installation time may increase.
|
||||
- Additional device storage may be required, especially for A/B layouts.
|
||||
- Individual components cannot easily maintain independent release cycles.
|
||||
- Hardware-specific differences may require multiple variants of the same product release.
|
||||
- Every image variant must be built, tested, signed, maintained, and assigned to the correct devices.
|
||||
|
||||
## Special Considerations
|
||||
|
||||
### Firmware
|
||||
|
||||
For many microcontroller-based devices, full-image replacement is the natural and often the only practical OTA model.
|
||||
|
||||
Such devices commonly do not provide:
|
||||
|
||||
- a filesystem;
|
||||
- a package manager;
|
||||
- independently deployable software components;
|
||||
- a practical mechanism for updating only part of the firmware.
|
||||
|
||||
In this context, OTA normally means replacing the complete firmware image.
|
||||
|
||||
### Full-System Testing
|
||||
|
||||
Full-system image updates do not reduce the required scope of testing.
|
||||
|
||||
Even when only one component changes, the complete resulting image must be tested because the change may introduce regressions elsewhere in the system.
|
||||
|
||||
The advantage is not that fewer tests are required. The advantage is that the exact image installed on the device is the same image that passed validation.
|
||||
|
||||
## Limitations
|
||||
|
||||
Full-system image updates are most effective when one software release corresponds to one clearly defined product version.
|
||||
|
||||
The approach becomes more difficult to maintain when the fleet contains many devices that are functionally equivalent but differ in small hardware-dependent details, such as:
|
||||
|
||||
- peripheral hardware revisions;
|
||||
- bundled peripheral firmware versions;
|
||||
- microcode state;
|
||||
- hardware-specific migration history;
|
||||
- other small platform-dependent differences.
|
||||
|
||||
These differences may create several variants of the same release.
|
||||
|
||||
Each variant must be:
|
||||
|
||||
- built;
|
||||
- regression-tested;
|
||||
- signed;
|
||||
- stored and maintained;
|
||||
- assigned to the correct group of devices.
|
||||
|
||||
If the number of variants grows quickly, the operational advantages of a single full-system image decrease while release and validation costs increase.
|
||||
|
||||
## Applicability
|
||||
|
||||
This approach is a strong choice when:
|
||||
|
||||
- the software or firmware is treated as one deployable product;
|
||||
- the cost of an inconsistent system state is high;
|
||||
- physical recovery is difficult or expensive;
|
||||
- reliable rollback is required;
|
||||
- the fleet is relatively homogeneous;
|
||||
- releases are centrally built and validated.
|
||||
|
||||
Other update strategies should be considered when:
|
||||
|
||||
- components have independent lifecycles;
|
||||
- the fleet contains many small hardware variants;
|
||||
- network usage is a primary constraint;
|
||||
- available storage is severely limited;
|
||||
- some components must be updated much more frequently than the rest of the system;
|
||||
- full regression testing of every image variant is too expensive.
|
||||
|
||||
## Alternatives Considered
|
||||
|
||||
The following alternatives are addressed in separate ADRs:
|
||||
|
||||
- package-based updates;
|
||||
- container-based updates;
|
||||
- versioned filesystem trees;
|
||||
- hybrid update strategies.
|
||||
136
adr/ADR-005-package-based-updates.md
Normal file
136
adr/ADR-005-package-based-updates.md
Normal file
@@ -0,0 +1,136 @@
|
||||
# ADR-005: Package-Based Updates
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
An OTA architecture must define the unit of update.
|
||||
|
||||
One possible approach is to update individual software components as independent packages.
|
||||
|
||||
Unlike full-system image updates, this approach treats the system image as a stable software platform on top of which application components evolve independently.
|
||||
|
||||
In embedded Linux, the platform is typically updated separately, while application components are distributed, maintained, and released as independent packages.
|
||||
|
||||
## Decision
|
||||
|
||||
Use package-based updates for components intentionally designed as independent deployment units that depend only on a specific version of the system platform.
|
||||
|
||||
Compatibility between components should be achieved by architecture rather than determined dynamically during installation.
|
||||
|
||||
## Decision Drivers
|
||||
|
||||
Package-based updates are preferred when:
|
||||
|
||||
- the system platform evolves independently from application components;
|
||||
- application components have their own release lifecycle;
|
||||
- individual functionality must be updated without releasing a new system image;
|
||||
- minimizing OTA traffic is desirable;
|
||||
- hardware variants can be represented by different sets of installed components;
|
||||
- components are designed to be independent from one another.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- Individual components can be updated independently.
|
||||
- Small changes do not require a new system image.
|
||||
- OTA download size is reduced.
|
||||
- A single system image can support multiple hardware models.
|
||||
- Device functionality can be defined by the installed package set.
|
||||
- User data and runtime state are naturally preserved because updates affect only the software component rather than the entire system.
|
||||
- Previous component versions can be retained for downgrade.
|
||||
- Package authenticity, integrity, and delivery are delegated to the package manager.
|
||||
|
||||
### Negative
|
||||
|
||||
- Component independence must be designed into the architecture from the beginning.
|
||||
- Every platform version requires its own package repository.
|
||||
- The same logical component may exist as multiple binaries built for different platform versions.
|
||||
- The build infrastructure must maintain multiple platform-specific repositories.
|
||||
- A package manager does not replace an OTA manager responsible for update policy.
|
||||
|
||||
## Special Considerations
|
||||
|
||||
### Stable System Platform
|
||||
|
||||
Package-based updates do not eliminate the system image.
|
||||
|
||||
The system platform is updated independently and defines:
|
||||
|
||||
- available libraries;
|
||||
- ABI compatibility;
|
||||
- core system services;
|
||||
- repository structure.
|
||||
|
||||
Each platform version uses its own repository containing packages built and validated specifically for that platform.
|
||||
|
||||
Compatibility between the platform and application components is therefore established by repository design rather than runtime dependency resolution.
|
||||
|
||||
### Component Independence
|
||||
|
||||
This approach depends on the application architecture.
|
||||
|
||||
Components should be designed to:
|
||||
|
||||
- update independently;
|
||||
- avoid coordinated updates with unrelated components;
|
||||
- depend only on the system platform;
|
||||
- expose stable interfaces.
|
||||
|
||||
Relationships between packages may describe installation order or device composition, but should not create a tightly coupled application.
|
||||
|
||||
### OTA Manager and Package Manager
|
||||
|
||||
The package manager is responsible for:
|
||||
|
||||
- downloading packages;
|
||||
- verifying signatures and integrity;
|
||||
- executing installation scripts;
|
||||
- replacing files;
|
||||
- retaining previously installed versions.
|
||||
|
||||
The OTA manager is responsible for update policy, including when updates are allowed, network restrictions, user approval, and operational conditions.
|
||||
|
||||
Using a package manager alone does not constitute a complete OTA solution.
|
||||
|
||||
### Component Rollback
|
||||
|
||||
Each component can be upgraded and downgraded independently.
|
||||
|
||||
Most package managers retain previously installed versions, allowing an individual component to be rolled back without reverting the entire system.
|
||||
|
||||
## Limitations
|
||||
|
||||
Package-based updates are most effective when software is intentionally designed as a collection of independent components.
|
||||
|
||||
If strong coupling develops between components or multiple packages frequently require synchronized updates, the benefits of this approach diminish and maintenance complexity increases.
|
||||
|
||||
## Applicability
|
||||
|
||||
This approach is well suited when:
|
||||
|
||||
- the system platform changes less frequently than application components;
|
||||
- device features evolve independently;
|
||||
- OTA bandwidth should be minimized;
|
||||
- individual fixes must be released quickly;
|
||||
- multiple hardware models share the same system platform.
|
||||
|
||||
Another approach should be considered when:
|
||||
|
||||
- the device is effectively a single firmware image;
|
||||
- the entire system must always be tested and released as one product;
|
||||
- most components change together;
|
||||
- inconsistent system states are unacceptable;
|
||||
- almost every update affects the complete software stack.
|
||||
|
||||
## Alternatives Considered
|
||||
|
||||
The following alternatives are described in separate ADRs:
|
||||
|
||||
- full-system image updates;
|
||||
- container-based updates;
|
||||
- versioned filesystem trees;
|
||||
- hybrid update strategies.
|
||||
177
adr/ADR-006-container-based-updates.md
Normal file
177
adr/ADR-006-container-based-updates.md
Normal file
@@ -0,0 +1,177 @@
|
||||
# ADR-006: Container-Based Updates
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
An OTA architecture must define the unit of update.
|
||||
|
||||
One possible approach is to distribute software as a container that includes an application together with the user-space environment required to run it.
|
||||
|
||||
Unlike a package-based update, a container is treated as a self-contained delivery unit that may include the application, libraries, runtime components, configuration, and other dependencies.
|
||||
|
||||
In embedded Linux, container-based delivery is usually less common than in server environments and should be introduced only when it solves a clear architectural or organizational problem.
|
||||
|
||||
## Decision
|
||||
|
||||
Use container-based updates only when an application or subsystem must be delivered and maintained as an independent software product with minimal dependence on the target platform's user-space environment.
|
||||
|
||||
Do not use containerization as a substitute for a well-designed system architecture.
|
||||
|
||||
## Decision Drivers
|
||||
|
||||
Container-based updates are preferred when:
|
||||
|
||||
- the software subsystem is owned and maintained by a different organization;
|
||||
- the supplier must support multiple compatible target platforms;
|
||||
- the application must be delivered with its own user-space environment;
|
||||
- changes to the host platform should have minimal impact on the delivered subsystem;
|
||||
- the application or subsystem is maintained as an independent software product;
|
||||
- the additional storage and runtime overhead are acceptable.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- The application is delivered together with its required user-space environment.
|
||||
- Dependence on the host system's user-space libraries is reduced.
|
||||
- A supplier can maintain a single delivery for multiple compatible platforms.
|
||||
- Independent subsystems can follow their own release lifecycle.
|
||||
- Conflicts between library versions used by separate subsystems are less likely.
|
||||
- The ownership boundary between the subsystem supplier and the device integrator becomes explicit.
|
||||
|
||||
### Negative
|
||||
|
||||
- The delivered artifact is larger.
|
||||
- Libraries and supporting components may be duplicated across containers.
|
||||
- A container update typically replaces the container as a whole.
|
||||
- Diagnosing and maintaining third-party containers is more difficult.
|
||||
- The device integrator loses visibility and control over part of the delivered software.
|
||||
- The container may become a black box whose internal composition is controlled by the supplier.
|
||||
- The integrator may not know which library versions, build options, patches, tools, or auxiliary processes are present.
|
||||
- Internal defects often cannot be corrected without supplier involvement.
|
||||
- The container runtime becomes an additional part of the maintained platform.
|
||||
|
||||
## Special Considerations
|
||||
|
||||
### Independent Supplier
|
||||
|
||||
Container-based delivery is most useful when a subsystem is supplied by a third party.
|
||||
|
||||
It allows the supplier to distribute an almost complete software environment without adapting it to every customer's user-space platform.
|
||||
|
||||
For the system integrator, this may simplify initial integration. It also reduces control over the subsystem's internal design and maintenance.
|
||||
|
||||
Delivery convenience does not remove operational responsibility. The integrator must still supervise startup, resource usage, logging, health, updates, and interaction with the rest of the device.
|
||||
|
||||
### Containers as Black Boxes
|
||||
|
||||
Bundling a complete user-space environment transfers part of the control from the integrator to the supplier.
|
||||
|
||||
Even when familiar libraries are included, the integrator may not be able to determine:
|
||||
|
||||
- how they were built;
|
||||
- which patches were applied;
|
||||
- which optional features were enabled;
|
||||
- which additional dependencies were included;
|
||||
- which background processes are started;
|
||||
- what changed between container releases.
|
||||
|
||||
This complicates auditing, debugging, security review, and impact analysis.
|
||||
|
||||
The less transparent the container contents are, the more the integrator depends on the supplier's release process, documentation, and long-term support.
|
||||
|
||||
### Containerization Does Not Replace Architecture
|
||||
|
||||
If the entire software stack is developed and maintained by one organization, the benefits of containerization are often limited.
|
||||
|
||||
A well-designed package-based architecture can frequently provide similar lifecycle independence without introducing a separate user-space environment for each component.
|
||||
|
||||
Containerization does not correct:
|
||||
|
||||
- tightly coupled components;
|
||||
- unstable interfaces;
|
||||
- incorrect state management;
|
||||
- poor service lifecycle design;
|
||||
- inadequate observability;
|
||||
- weak diagnostics.
|
||||
|
||||
Wrapping a poorly designed component in a container does not make it independent.
|
||||
|
||||
### Fault Isolation
|
||||
|
||||
Containers are often presented as a fault-tolerance mechanism.
|
||||
|
||||
In most embedded Linux systems, user-space processes are already isolated by the operating system. A correctly designed service can fail without terminating unrelated services and can be restarted by a supervisor, watchdog, or service manager.
|
||||
|
||||
Restarting a container therefore does not necessarily provide stronger fault isolation than restarting a process, and may consume more time and resources.
|
||||
|
||||
Containerization also does not protect against a shared kernel failure. A kernel panic remains fatal to both containerized and non-containerized user-space architectures.
|
||||
|
||||
Systems that isolate complete operating systems under a hypervisor or a safety-certified RTOS use a different architectural model and are outside the scope of this ADR.
|
||||
|
||||
### Different Library Versions
|
||||
|
||||
Containers can be useful when independent subsystems require incompatible versions of the same libraries.
|
||||
|
||||
However, this does not always require full containerization.
|
||||
|
||||
In many embedded Linux systems, a component can ship with private libraries in its own directory and load them independently from the platform-wide versions.
|
||||
|
||||
Different library requirements alone are therefore not sufficient justification for introducing containers.
|
||||
|
||||
## Limitations
|
||||
|
||||
Containerization does not eliminate software maintenance.
|
||||
|
||||
It transfers part of the responsibility for the user-space environment from the device developer to the container supplier.
|
||||
|
||||
If the supplier does not provide timely bug fixes, security updates, or new compatible releases, the integrator may have little practical ability to maintain the container internals independently.
|
||||
|
||||
Container updates and rollbacks may also replace the entire bundled user space even when only one internal component changed.
|
||||
|
||||
This approach therefore requires:
|
||||
|
||||
- a high level of trust in the supplier;
|
||||
- a transparent release process;
|
||||
- documented container contents;
|
||||
- long-term maintenance commitments;
|
||||
- clear diagnostic and update procedures.
|
||||
|
||||
## Applicability
|
||||
|
||||
This approach is well suited when:
|
||||
|
||||
- the subsystem is supplied and maintained by an independent organization;
|
||||
- the supplier must minimize dependence on the customer's user-space platform;
|
||||
- the application or subsystem is an independently maintained product;
|
||||
- a clear responsibility boundary between supplier and integrator is required;
|
||||
- additional storage, memory, and runtime overhead are acceptable;
|
||||
- independent delivery is more valuable than full control over the implementation.
|
||||
|
||||
Another approach should be considered when:
|
||||
|
||||
- the complete software stack is developed by one organization;
|
||||
- components can be delivered as independent packages;
|
||||
- storage, memory, or OTA download size are critical;
|
||||
- full control over all device software is required;
|
||||
- simple component-level rollback is important;
|
||||
- containers are used only as packaging for ordinary services;
|
||||
- expected fault tolerance is already provided by Linux process isolation and supervisor mechanisms.
|
||||
|
||||
## Alternatives Considered
|
||||
|
||||
The following alternatives are described in separate ADRs:
|
||||
|
||||
- full-system image updates;
|
||||
- package-based updates;
|
||||
- versioned filesystem trees;
|
||||
- hybrid update strategies.
|
||||
|
||||
## Summary
|
||||
|
||||
Containerization primarily solves organizational problems related to the delivery and maintenance of independently owned software subsystems.
|
||||
|
||||
When all software components are controlled by one engineering organization, a well-designed package-based architecture often provides the same practical benefits with significantly less system complexity.
|
||||
83
adr/readme.md
Normal file
83
adr/readme.md
Normal file
@@ -0,0 +1,83 @@
|
||||
# Architecture Decision Records
|
||||
|
||||
This directory contains Architecture Decision Records (ADRs) used throughout the OTA Reference Design.
|
||||
|
||||
Unlike ADRs maintained within a single commercial project, these documents are **not** a historical record of decisions made for one specific product.
|
||||
|
||||
Instead, they serve as **reference architectural decisions** illustrating how an experienced engineer might reason about common OTA design problems under different technical constraints.
|
||||
|
||||
## Purpose
|
||||
|
||||
Each ADR documents:
|
||||
|
||||
- the engineering context;
|
||||
- the decision being considered;
|
||||
- the factors driving that decision;
|
||||
- the expected consequences;
|
||||
- the situations where the decision is appropriate;
|
||||
- the situations where another approach may be preferable.
|
||||
|
||||
The goal is not to declare one solution universally correct.
|
||||
|
||||
The goal is to explain **why** a particular decision would be reasonable for a particular class of systems.
|
||||
|
||||
## Relationship to the Documentation
|
||||
|
||||
The repository intentionally separates three different types of documents.
|
||||
|
||||
### Concept Articles (`docs/`)
|
||||
|
||||
Concept articles explain the design space.
|
||||
|
||||
They answer questions such as:
|
||||
|
||||
- What approaches exist?
|
||||
- How do they work?
|
||||
- What problems do they solve?
|
||||
- What trade-offs do they introduce?
|
||||
|
||||
Their purpose is education.
|
||||
|
||||
### Architecture Decision Records (`adr/`)
|
||||
|
||||
ADRs answer a different question:
|
||||
|
||||
> Given a particular set of engineering constraints, what decision would be made, and why?
|
||||
|
||||
They intentionally focus on architectural reasoning rather than implementation details.
|
||||
|
||||
### Decision Matrix
|
||||
|
||||
After the individual ADRs, the series concludes with a decision matrix that compares the approaches and helps relate system constraints to architectural choices.
|
||||
|
||||
The matrix is not intended to automatically select a solution. Instead, it provides a structured way to evaluate trade-offs.
|
||||
|
||||
## Educational Nature
|
||||
|
||||
These ADRs are examples.
|
||||
|
||||
A real product may reach different conclusions depending on its:
|
||||
|
||||
- reliability requirements;
|
||||
- hardware architecture;
|
||||
- network constraints;
|
||||
- operational model;
|
||||
- regulatory requirements;
|
||||
- maintenance strategy;
|
||||
- business priorities.
|
||||
|
||||
For that reason, every ADR should be read as:
|
||||
|
||||
> "Given these assumptions, this is the architectural decision we would make."
|
||||
|
||||
rather than:
|
||||
|
||||
> "This is the only correct solution."
|
||||
|
||||
## Design Philosophy
|
||||
|
||||
Throughout this repository, architectural decisions follow the same principle:
|
||||
|
||||
> **Architecture should be driven by system constraints, not by technology preferences.**
|
||||
|
||||
The objective is to teach the engineering process behind OTA design rather than promote a specific framework, product, or update mechanism.
|
||||
38
adr/template.md
Normal file
38
adr/template.md
Normal file
@@ -0,0 +1,38 @@
|
||||
# ADR-NNNN: Decision Title
|
||||
|
||||
* **Status:** Proposed
|
||||
* **Date:** YYYY-MM-DD
|
||||
|
||||
## Context
|
||||
|
||||
Describe the problem or architectural question that requires a decision.
|
||||
|
||||
Explain the relevant requirements, constraints, risks, and assumptions. Keep this section focused on why the decision is needed.
|
||||
|
||||
## Decision
|
||||
|
||||
Describe the selected approach clearly and directly.
|
||||
|
||||
## Alternatives Considered
|
||||
|
||||
### Alternative 1
|
||||
|
||||
Briefly describe the alternative and why it was not selected.
|
||||
|
||||
### Alternative 2
|
||||
|
||||
Briefly describe the alternative and why it was not selected.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
* List the main benefits of the decision.
|
||||
|
||||
### Negative
|
||||
|
||||
* List the costs, limitations, and trade-offs introduced by the decision.
|
||||
|
||||
## Notes
|
||||
|
||||
Add implementation details, unresolved questions, or links to related documents when necessary.
|
||||
262
docs/00-series-overview.md
Normal file
262
docs/00-series-overview.md
Normal file
@@ -0,0 +1,262 @@
|
||||
# OTA Reference Design
|
||||
|
||||
> A practical guide to designing reliable, secure, and maintainable over-the-air (OTA) update systems for embedded Linux devices.
|
||||
|
||||
---
|
||||
|
||||
## About This Repository
|
||||
|
||||
This repository presents a practical reference design for building over-the-air (OTA) update systems for embedded Linux devices.
|
||||
|
||||
Rather than documenting a particular framework or vendor-specific solution, it focuses on the engineering principles that make OTA systems reliable, secure, and maintainable.
|
||||
|
||||
The goal is to explain **why** modern OTA systems are designed the way they are, what problems they solve, and what trade-offs different approaches involve.
|
||||
|
||||
Although many examples use Raspberry Pi as a demonstration platform, the concepts are applicable to a wide range of embedded Linux systems.
|
||||
|
||||
---
|
||||
|
||||
## Why This Repository Exists
|
||||
|
||||
There is no shortage of documentation for OTA frameworks.
|
||||
|
||||
You can easily find documentation for tools such as:
|
||||
|
||||
- Mender
|
||||
- RAUC
|
||||
- SWUpdate
|
||||
- OSTree
|
||||
- A/B Updates
|
||||
|
||||
However, these resources usually explain **how to use a particular tool**, not **why OTA systems are designed this way**.
|
||||
|
||||
Questions like these are often left unanswered:
|
||||
|
||||
- Why do many devices use A/B partitions?
|
||||
- Why are bootloaders involved in the update process?
|
||||
- Why is rollback necessary?
|
||||
- What happens if power is lost during an update?
|
||||
- Why are update images signed?
|
||||
- Why are atomic updates important?
|
||||
- How do production devices remain recoverable after failures?
|
||||
|
||||
This repository attempts to answer those engineering questions.
|
||||
|
||||
---
|
||||
|
||||
## What You'll Learn
|
||||
|
||||
Throughout this series we will explore topics including:
|
||||
|
||||
- OTA architecture
|
||||
- Boot process
|
||||
- Update strategies
|
||||
- Full-image vs package updates
|
||||
- A/B partition layouts
|
||||
- Bootloader interaction
|
||||
- Rollback mechanisms
|
||||
- Atomic updates
|
||||
- Image verification
|
||||
- Digital signatures
|
||||
- Secure Boot
|
||||
- Failure recovery
|
||||
- Delta updates
|
||||
- Version management
|
||||
- Testing strategies
|
||||
- Production deployment considerations
|
||||
|
||||
The emphasis is always on understanding the underlying design rather than memorizing a particular implementation.
|
||||
|
||||
---
|
||||
|
||||
## Repository Structure
|
||||
|
||||
```text
|
||||
ota-reference-design/
|
||||
│
|
||||
├── docs/
|
||||
│ ├── 00-series-overview.md
|
||||
│ ├── 01-ota-introduction.md
|
||||
│ ├── 02-update-strategies.md
|
||||
│ ├── ...
|
||||
│
|
||||
├── diagrams/
|
||||
│ └── plantuml/
|
||||
│
|
||||
├── examples/
|
||||
│ ├── raspberry-pi/
|
||||
│ ├── qemu/
|
||||
│ └── simulations/
|
||||
│
|
||||
├── adr/
|
||||
│ ├── ADR-001-repository-name.md
|
||||
│ ├── ADR-002-license.md
|
||||
│ └── ADR-003-reference-design.md
|
||||
│
|
||||
└── README.md
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Learning Path
|
||||
|
||||
The chapters are designed to build on each other.
|
||||
|
||||
```text
|
||||
Introduction
|
||||
│
|
||||
▼
|
||||
Update Strategies
|
||||
│
|
||||
▼
|
||||
System Architecture
|
||||
│
|
||||
▼
|
||||
Boot Process
|
||||
│
|
||||
▼
|
||||
Storage Layout
|
||||
│
|
||||
▼
|
||||
Atomic Updates
|
||||
│
|
||||
▼
|
||||
Rollback
|
||||
│
|
||||
▼
|
||||
Security
|
||||
│
|
||||
▼
|
||||
Testing
|
||||
│
|
||||
▼
|
||||
Production Deployment
|
||||
```
|
||||
|
||||
While each chapter can be read independently, following the series in order provides a much deeper understanding of the complete system.
|
||||
|
||||
---
|
||||
|
||||
## Engineering Philosophy
|
||||
|
||||
This repository is intentionally different from product documentation.
|
||||
|
||||
Instead of presenting a single "correct" solution, every topic discusses:
|
||||
|
||||
- why a particular design exists;
|
||||
- which problem it solves;
|
||||
- what alternatives are available;
|
||||
- what trade-offs each approach introduces;
|
||||
- when another solution may be more appropriate.
|
||||
|
||||
Real-world engineering is rarely about choosing the only correct answer.
|
||||
|
||||
It is about understanding constraints and making informed decisions.
|
||||
|
||||
---
|
||||
|
||||
## The Reference Design
|
||||
|
||||
The architecture presented throughout this repository is a coherent reference design.
|
||||
|
||||
Real products may use different technologies or frameworks while following the same architectural principles.
|
||||
|
||||
For example, one project may use RAUC, another SWUpdate, and another a completely custom implementation.
|
||||
|
||||
The implementation details may differ.
|
||||
|
||||
The underlying engineering principles usually do not.
|
||||
|
||||
---
|
||||
|
||||
## Practical Examples
|
||||
|
||||
Where possible, theoretical discussions are accompanied by practical material, including:
|
||||
|
||||
- architecture diagrams
|
||||
- boot sequence walkthroughs
|
||||
- storage layout examples
|
||||
- failure scenarios
|
||||
- Raspberry Pi demonstrations
|
||||
- QEMU-based experiments
|
||||
- implementation notes
|
||||
|
||||
The objective is to connect high-level architecture with practical implementation.
|
||||
|
||||
---
|
||||
|
||||
## Intended Audience
|
||||
|
||||
This repository is intended for:
|
||||
|
||||
- Embedded Linux developers
|
||||
- Firmware engineers
|
||||
- Embedded software engineers
|
||||
- System architects
|
||||
- Students learning embedded systems
|
||||
- Engineers preparing for technical interviews
|
||||
- Anyone interested in understanding OTA system design
|
||||
|
||||
No prior experience with OTA frameworks is assumed.
|
||||
|
||||
---
|
||||
|
||||
## How to Read This Repository
|
||||
|
||||
If you are new to OTA systems, simply start with Chapter 1 and continue in order.
|
||||
|
||||
If you already have embedded Linux experience, feel free to jump directly to topics that interest you.
|
||||
|
||||
If you are looking for implementation details, the accompanying examples provide practical demonstrations of the concepts discussed in the documentation.
|
||||
|
||||
---
|
||||
|
||||
## OTA at a Glance
|
||||
|
||||
The following diagram illustrates the overall update lifecycle that will be explored throughout this repository.
|
||||
|
||||
```text
|
||||
OTA Server
|
||||
│
|
||||
Signed Update
|
||||
│
|
||||
┌───────────▼───────────┐
|
||||
│ Download Manager │
|
||||
└───────────┬───────────┘
|
||||
│
|
||||
Verify Signature
|
||||
│
|
||||
Verify Integrity
|
||||
│
|
||||
Install Update
|
||||
│
|
||||
Mark Boot Target
|
||||
│
|
||||
Reboot Device
|
||||
│
|
||||
Bootloader Decision
|
||||
┌─────────┴─────────┐
|
||||
│ │
|
||||
Boot Success Boot Failure
|
||||
│ │
|
||||
▼ ▼
|
||||
Commit Update Rollback
|
||||
```
|
||||
|
||||
Each stage of this process will be examined in detail in the chapters that follow.
|
||||
|
||||
---
|
||||
|
||||
## Contributing
|
||||
|
||||
Contributions, suggestions, and discussions are welcome.
|
||||
|
||||
If you have ideas for improvements, additional examples, or alternative approaches, feel free to open an issue or submit a pull request.
|
||||
|
||||
---
|
||||
|
||||
## License
|
||||
|
||||
This repository is released under the MIT License.
|
||||
|
||||
See the LICENSE file for details.
|
||||
227
docs/01-what-is-an-ota-update.md
Normal file
227
docs/01-what-is-an-ota-update.md
Normal file
@@ -0,0 +1,227 @@
|
||||
# What Is an OTA Update?
|
||||
|
||||
> **Series:** OTA Reference Design
|
||||
>
|
||||
> This article is the first chapter of a practical reference design describing how reliable and secure over-the-air software updates are built for embedded Linux devices.
|
||||
|
||||
---
|
||||
|
||||
# Introduction
|
||||
|
||||
Almost every modern connected device receives software updates remotely.
|
||||
|
||||
Phones do it.
|
||||
|
||||
Cars do it.
|
||||
|
||||
Industrial controllers do it.
|
||||
|
||||
Medical devices do it.
|
||||
|
||||
Consumer electronics quietly update themselves while nobody is watching.
|
||||
|
||||
This process is commonly known as an **Over-the-Air (OTA) update**.
|
||||
|
||||
At first glance, the idea seems simple:
|
||||
|
||||
> Download new software and install it.
|
||||
|
||||
In reality, OTA is one of the most challenging reliability problems in embedded systems.
|
||||
|
||||
---
|
||||
|
||||
# Why OTA Is Different
|
||||
|
||||
Updating software on a desktop computer is usually forgiving.
|
||||
|
||||
If something goes wrong, the user can often retry the installation, download the package again, or reinstall the operating system.
|
||||
|
||||
Embedded devices rarely have that luxury.
|
||||
|
||||
Imagine a device installed:
|
||||
|
||||
- on the roof of a building;
|
||||
- inside industrial equipment;
|
||||
- on a remote oil pipeline;
|
||||
- in a laboratory instrument;
|
||||
- in an autonomous vehicle.
|
||||
|
||||
A failed update may leave the device completely unreachable.
|
||||
|
||||
Nobody may be available to reconnect a keyboard, attach a monitor, or reflash storage.
|
||||
|
||||
For embedded systems, software updates must be designed with failure as an expected condition rather than an exceptional one.
|
||||
|
||||
---
|
||||
|
||||
# The Real Problem
|
||||
|
||||
The primary goal of an OTA system is surprisingly simple:
|
||||
|
||||
> **Replace the software while always preserving a path to recovery.**
|
||||
|
||||
Everything else exists to support this objective.
|
||||
|
||||
Notice that this definition says nothing about how the update is delivered or installed.
|
||||
|
||||
The engineering problem remains the same regardless of the implementation.
|
||||
|
||||
---
|
||||
|
||||
# Many Ways to Solve the Same Problem
|
||||
|
||||
Different products solve OTA updates in different ways.
|
||||
|
||||
An update may be based on:
|
||||
|
||||
- complete system images;
|
||||
- software packages;
|
||||
- application bundles;
|
||||
- containers;
|
||||
- custom update formats.
|
||||
|
||||
These are implementation choices.
|
||||
|
||||
Each approach has its own strengths, weaknesses, and trade-offs.
|
||||
|
||||
Throughout this series we will explore these options and discuss where each of them makes sense.
|
||||
|
||||
---
|
||||
|
||||
# Typical Failure Scenarios
|
||||
|
||||
A robust OTA implementation assumes that failures are inevitable.
|
||||
|
||||
Examples include:
|
||||
|
||||
- power loss during installation;
|
||||
- interrupted network connection;
|
||||
- corrupted download;
|
||||
- damaged storage;
|
||||
- software crash during the first boot;
|
||||
- incompatible configuration;
|
||||
- interrupted filesystem writes;
|
||||
- unexpected reboot.
|
||||
|
||||
None of these situations is unusual.
|
||||
|
||||
If enough devices are deployed, every one of them will eventually happen.
|
||||
|
||||
The question is never **if**.
|
||||
|
||||
Only **when**.
|
||||
|
||||
---
|
||||
|
||||
# OTA Is a System, Not a Feature
|
||||
|
||||
OTA is often imagined as a single application responsible for installing updates.
|
||||
|
||||
In practice, it is an entire system composed of multiple cooperating components.
|
||||
|
||||
A typical embedded Linux solution may include:
|
||||
|
||||
- bootloader;
|
||||
- Linux kernel;
|
||||
- root filesystem;
|
||||
- update agent;
|
||||
- storage layout;
|
||||
- cryptographic verification;
|
||||
- backend services;
|
||||
- device identity;
|
||||
- rollback mechanism;
|
||||
- health monitoring.
|
||||
|
||||
Each component has a specific responsibility.
|
||||
|
||||
Only together do they provide a reliable update process.
|
||||
|
||||
---
|
||||
|
||||
# A Better Mental Model
|
||||
|
||||
Instead of thinking:
|
||||
|
||||
```
|
||||
download
|
||||
↓
|
||||
install
|
||||
```
|
||||
|
||||
think:
|
||||
|
||||
```
|
||||
prepare
|
||||
↓
|
||||
verify
|
||||
↓
|
||||
store safely
|
||||
↓
|
||||
activate
|
||||
↓
|
||||
boot
|
||||
↓
|
||||
health check
|
||||
↓
|
||||
commit
|
||||
│
|
||||
└── rollback if necessary
|
||||
```
|
||||
|
||||
Almost every production OTA solution follows some variation of this workflow.
|
||||
|
||||
The individual technologies may differ.
|
||||
|
||||
The underlying principles remain remarkably similar.
|
||||
|
||||
---
|
||||
|
||||
# What This Series Covers
|
||||
|
||||
Rather than focusing on a particular framework or vendor, this repository explains the engineering principles behind reliable OTA systems.
|
||||
|
||||
Topics include:
|
||||
|
||||
- OTA architectures;
|
||||
- update strategies;
|
||||
- boot process;
|
||||
- storage layouts;
|
||||
- A/B partitioning;
|
||||
- rollback mechanisms;
|
||||
- image verification;
|
||||
- cryptographic signatures;
|
||||
- update servers;
|
||||
- recovery strategies;
|
||||
- production considerations.
|
||||
|
||||
Examples will use embedded Linux running on Raspberry Pi, but the concepts apply to many embedded platforms.
|
||||
|
||||
---
|
||||
|
||||
# Summary
|
||||
|
||||
OTA updates are often described as "remote software updates."
|
||||
|
||||
While technically correct, this definition misses the real engineering challenge.
|
||||
|
||||
The true objective is ensuring that **the device remains recoverable after every possible failure during the update process.**
|
||||
|
||||
Everything else—from storage layouts to cryptographic signatures—exists to support that goal.
|
||||
|
||||
---
|
||||
|
||||
## Key Takeaways
|
||||
|
||||
- OTA is fundamentally a reliability problem.
|
||||
- Failures must be expected, not treated as exceptions.
|
||||
- Multiple implementation strategies exist for OTA systems.
|
||||
- The core objective is always safe recovery.
|
||||
- Technologies change, but the engineering principles remain the same.
|
||||
|
||||
---
|
||||
|
||||
## Next Article
|
||||
|
||||
The next chapter explores the first major architectural decision in any OTA system:
|
||||
|
||||
> **Update Strategies: Full Images, Packages, or Something Else?**
|
||||
864
docs/02-update_strategies.md
Normal file
864
docs/02-update_strategies.md
Normal file
@@ -0,0 +1,864 @@
|
||||
# 02. Choosing an OTA Update Strategy: Full Images, Packages, or Something Else?
|
||||
|
||||
In the previous chapter, we defined an OTA update as the controlled delivery and application of a new version of software or data to a remote device.
|
||||
|
||||
The next natural question is:
|
||||
|
||||
> How should such an update actually be performed?
|
||||
|
||||
At first glance, the answer seems simple. There are full-image updates, packages, binary deltas, containers, atomic filesystem trees, and other well-known approaches. All that remains is to compare them and choose the best one.
|
||||
|
||||
But this framing starts too late.
|
||||
|
||||
Before choosing a technology, we need to answer a more fundamental question:
|
||||
|
||||
> **What exactly are we trying to update?**
|
||||
|
||||
The entire system? The operating system? A single application? A set of files? Configuration? Certificates? Maps? A machine-learning model?
|
||||
|
||||
The answer determines almost everything else: the update unit, storage requirements, verification method, activation mechanism, rollback capability, and acceptable downtime.
|
||||
|
||||
In addition, some commonly compared approaches describe different aspects of the process. A full image or a package defines **what the update unit is**. An A/B layout defines **where the new version is prepared and how the system switches between versions**. A binary delta defines **how the amount of transferred data is reduced**. A cryptographic signature defines **how the origin of the update is verified**.
|
||||
|
||||
These are not necessarily mutually exclusive choices.
|
||||
|
||||
For example, a device may use:
|
||||
|
||||
- full system images as the update unit;
|
||||
- an A/B layout for safe activation;
|
||||
- binary deltas to reduce network traffic;
|
||||
- cryptographic signatures to verify authenticity;
|
||||
- a separate mechanism for configuration updates.
|
||||
|
||||
Therefore, choosing an OTA strategy is not selecting one item from a short list. It is the design of several related mechanisms.
|
||||
|
||||
---
|
||||
|
||||
## 1. What Exactly Is Being Updated?
|
||||
|
||||
OTA is often associated exclusively with device firmware. For a small microcontroller, this may be a sufficiently accurate description: new firmware may indeed replace almost all executable content on the device.
|
||||
|
||||
An embedded Linux system is usually more complex.
|
||||
|
||||
It may include:
|
||||
|
||||
```text
|
||||
Bootloader
|
||||
Kernel
|
||||
Device tree
|
||||
Root filesystem
|
||||
System libraries
|
||||
System services
|
||||
Applications
|
||||
Configuration
|
||||
Persistent data
|
||||
Certificates
|
||||
Content
|
||||
```
|
||||
|
||||
These components have different lifecycles.
|
||||
|
||||
The kernel may be updated rarely. A user application may be updated every month. Configuration may change several times a day. A certificate may be replaced only before it expires. Maps or machine-learning models may have their own independent release cycle.
|
||||
|
||||
The first step is therefore to identify the possible update targets.
|
||||
|
||||
### 1.1. The Complete System
|
||||
|
||||
In this case, an update is treated as a transition from one complete device version to another:
|
||||
|
||||
```text
|
||||
System version N
|
||||
|
|
||||
v
|
||||
System version N + 1
|
||||
```
|
||||
|
||||
The new version may include the kernel, root filesystem, system libraries, services, and applications.
|
||||
|
||||
The defining property of this approach is that the system is delivered and tested as a whole. The device does not assemble the new version itself from many independently changing components. It receives a predefined state.
|
||||
|
||||
This improves reproducibility:
|
||||
|
||||
> The system version installed on the device should match the version built and tested by the build system.
|
||||
|
||||
This is the level at which full-image updates, A/B layouts, recovery partitions, and atomic switching between versions are usually considered.
|
||||
|
||||
### 1.2. The Operating System
|
||||
|
||||
Sometimes the system layer must be updated without replacing the entire contents of the device.
|
||||
|
||||
Update targets may include:
|
||||
|
||||
- the kernel;
|
||||
- system libraries;
|
||||
- system services;
|
||||
- drivers and modules;
|
||||
- individual distribution packages.
|
||||
|
||||
This approach often relies on a package manager and the existing infrastructure of a Linux distribution.
|
||||
|
||||
In this case, the device does not receive a completely prepared new state. It transforms the current system into a new one by installing, removing, or replacing individual components.
|
||||
|
||||
### 1.3. Applications
|
||||
|
||||
In many products, the base platform changes rarely while application software has a much shorter release cycle.
|
||||
|
||||
For example:
|
||||
|
||||
```text
|
||||
Embedded Linux
|
||||
├── Device manager
|
||||
├── Communication service
|
||||
├── Web interface
|
||||
├── User interface
|
||||
└── Diagnostics
|
||||
```
|
||||
|
||||
An update may affect only one service or application.
|
||||
|
||||
This avoids replacing the entire system for a small application-level change. However, it introduces a new question: is the new application version compatible with the installed libraries, configuration, data schema, and other services?
|
||||
|
||||
The smaller the update unit, the smaller the change, but the larger the number of version combinations that may need to be supported.
|
||||
|
||||
### 1.4. Data and Content
|
||||
|
||||
OTA can update more than executable code.
|
||||
|
||||
Examples include:
|
||||
|
||||
- configuration;
|
||||
- calibration data;
|
||||
- certificates and key material;
|
||||
- maps;
|
||||
- voice packages;
|
||||
- dictionaries;
|
||||
- machine-learning models;
|
||||
- filtering rules;
|
||||
- static user-interface resources;
|
||||
- parameter databases or diagnostic-code databases.
|
||||
|
||||
Such data often cannot reasonably be included in a full system image every time it changes. It may require its own versioning, verification, compatibility, and rollback rules.
|
||||
|
||||
For example, a new machine-learning model may require a specific runtime version. A new configuration may be incompatible with an older application. A new certificate may require coordinated changes on both the client and server sides.
|
||||
|
||||
Therefore, even a data update remains an architectural problem rather than a simple file transfer.
|
||||
|
||||
### 1.5. Multiple Levels at the Same Time
|
||||
|
||||
A real product usually has more than one update unit.
|
||||
|
||||
For example:
|
||||
|
||||
```text
|
||||
Operating system -> full system image
|
||||
Applications -> packages or containers
|
||||
Maps and ML models -> separate artifacts
|
||||
Configuration -> small atomic documents
|
||||
Certificates -> dedicated rotation mechanism
|
||||
```
|
||||
|
||||
The operating system may be updated several times per year, applications monthly, data weekly, and configuration whenever necessary.
|
||||
|
||||
The question is therefore not:
|
||||
|
||||
> How is the device updated?
|
||||
|
||||
It is:
|
||||
|
||||
> **Which parts of the device are updated, how independently, and how often?**
|
||||
|
||||
---
|
||||
|
||||
## 2. What Decisions Make Up an OTA Strategy?
|
||||
|
||||
The term *update strategy* is often used too broadly. As a result, full images, A/B layouts, deltas, containers, and signatures appear in the same list even though they solve different problems.
|
||||
|
||||
It is useful to separate an OTA design into several independent questions.
|
||||
|
||||
### 2.1. Update Unit
|
||||
|
||||
What is the installable artifact?
|
||||
|
||||
- a full image;
|
||||
- a filesystem;
|
||||
- a package;
|
||||
- a container;
|
||||
- a data set.
|
||||
|
||||
### 2.2. Delivery Method
|
||||
|
||||
How is the artifact transferred to the device?
|
||||
|
||||
- in full;
|
||||
- compressed;
|
||||
- as a binary delta;
|
||||
- in blocks;
|
||||
- as a stream;
|
||||
- through a local gateway or intermediate cache.
|
||||
|
||||
### 2.3. Installation Method
|
||||
|
||||
How is the new version prepared?
|
||||
|
||||
- written to an inactive partition;
|
||||
- applied directly to the running filesystem;
|
||||
- assembled as a new filesystem tree;
|
||||
- unpacked into a separate directory;
|
||||
- imported as a new container image.
|
||||
|
||||
### 2.4. Activation Method
|
||||
|
||||
When does the new version become active?
|
||||
|
||||
- immediately after a file is replaced;
|
||||
- after a service restart;
|
||||
- after a device reboot;
|
||||
- after an atomic symlink or partition switch;
|
||||
- after confirmation by an external system.
|
||||
|
||||
### 2.5. Recovery Method
|
||||
|
||||
What happens if the update fails?
|
||||
|
||||
- return to the previous partition;
|
||||
- boot into a recovery system;
|
||||
- roll back a transaction;
|
||||
- reinstall the previous package;
|
||||
- return to the previous filesystem tree;
|
||||
- require manual service.
|
||||
|
||||
These decisions are related, but they are not the same.
|
||||
|
||||
For example, a full-image update may or may not use an A/B layout. A package update may be transferred in full or as a delta. A container may be only the application delivery mechanism, while the base operating system is updated with a separate system image.
|
||||
|
||||
---
|
||||
|
||||
## 3. Full-Image Updates
|
||||
|
||||
With a full-image update, the device receives a prebuilt image of the system or one of its major partitions.
|
||||
|
||||
The process usually looks like this:
|
||||
|
||||
```text
|
||||
Build system
|
||||
|
|
||||
| produces tested image
|
||||
v
|
||||
Update server
|
||||
|
|
||||
| delivers image
|
||||
v
|
||||
Device
|
||||
|
|
||||
| writes image
|
||||
v
|
||||
New system version
|
||||
```
|
||||
|
||||
Instead of changing individual components sequentially, the device receives a complete version.
|
||||
|
||||
### 3.1. Primary Advantage
|
||||
|
||||
The main advantage is a reproducible state.
|
||||
|
||||
If the image was built and tested as a whole, the device should end up in exactly that state after a successful installation.
|
||||
|
||||
With a correct implementation, this reduces the number of possible version combinations:
|
||||
|
||||
```text
|
||||
Version A
|
||||
Version B
|
||||
Version C
|
||||
```
|
||||
|
||||
instead of:
|
||||
|
||||
```text
|
||||
Kernel A + library B + service C + application D
|
||||
Kernel A + library C + service C + application D
|
||||
Kernel B + library C + service D + application D
|
||||
...
|
||||
```
|
||||
|
||||
The more independently updated components a system has, the faster the number of possible states grows.
|
||||
|
||||
### 3.2. Advantages
|
||||
|
||||
- predictable and reproducible system state;
|
||||
- convenient system-level testing;
|
||||
- no dependency resolution on the device;
|
||||
- natural compatibility with an A/B layout;
|
||||
- relatively clear rollback mechanism;
|
||||
- the system partition can remain immutable during normal operation;
|
||||
- easier identification of the exact installed component set.
|
||||
|
||||
### 3.3. Disadvantages
|
||||
|
||||
- large artifact size;
|
||||
- increased storage requirements;
|
||||
- long downloads over slow connections;
|
||||
- transfer of unchanged data;
|
||||
- user data and configuration require separate handling;
|
||||
- a small application change may require releasing a new system image.
|
||||
|
||||
### 3.4. Typical Uses
|
||||
|
||||
This approach is especially natural for appliance-like devices where the entire system is treated as one product:
|
||||
|
||||
- industrial controllers;
|
||||
- network devices;
|
||||
- automotive control units;
|
||||
- medical devices;
|
||||
- terminals;
|
||||
- specialized embedded Linux systems.
|
||||
|
||||
It is particularly attractive where reliability and reproducibility matter more than minimizing update size.
|
||||
|
||||
---
|
||||
|
||||
## 4. Package-Based Updates
|
||||
|
||||
With a package-based approach, the system is updated by installing or replacing individual packages.
|
||||
|
||||
These may use standard Linux mechanisms such as:
|
||||
|
||||
- `deb`;
|
||||
- `rpm`;
|
||||
- `opkg`;
|
||||
|
||||
or a custom package format created for a specific product.
|
||||
|
||||
A package usually contains:
|
||||
|
||||
- files;
|
||||
- metadata;
|
||||
- a version;
|
||||
- dependencies;
|
||||
- installation instructions;
|
||||
- sometimes pre-installation and post-installation scripts.
|
||||
|
||||
### 4.1. Primary Advantage
|
||||
|
||||
Packages allow only the changed components to be updated.
|
||||
|
||||
If one service has been fixed, there is no need to transfer and replace the entire filesystem.
|
||||
|
||||
This is especially convenient when the device is already based on a maintained Linux distribution and its components have relatively independent lifecycles.
|
||||
|
||||
### 4.2. Advantages
|
||||
|
||||
- smaller update size;
|
||||
- reuse of an existing ecosystem;
|
||||
- independent component release cycles;
|
||||
- only required changes need to be installed;
|
||||
- a familiar model for Linux engineers;
|
||||
- convenient dependency management when the repository is tightly controlled.
|
||||
|
||||
### 4.3. Disadvantages
|
||||
|
||||
- a large number of possible system states;
|
||||
- dependence on correct dependency resolution;
|
||||
- installation scripts may fail after partial execution;
|
||||
- system-wide atomicity is more difficult;
|
||||
- rolling back a package does not always restore the previous state;
|
||||
- data and configuration migrations may be irreversible;
|
||||
- the result may depend on the initial state of the device.
|
||||
|
||||
The last point is especially important.
|
||||
|
||||
If two devices begin the update from different states, the same sequence of package operations does not necessarily produce exactly the same final state.
|
||||
|
||||
### 4.4. A Package Manager Does Not Exclude OTA
|
||||
|
||||
Using a package manager does not make a mechanism “not real OTA.”
|
||||
|
||||
OTA describes the remote delivery and controlled application of an update. A package manager can absolutely be part of such a system.
|
||||
|
||||
The critical questions are not whether the implementation uses `apt`, `rpm`, `opkg`, or a custom installer, but:
|
||||
|
||||
- who controls the repository;
|
||||
- how package authenticity is verified;
|
||||
- which version transitions are allowed;
|
||||
- what happens if power is lost;
|
||||
- how partial installation is detected;
|
||||
- whether a working state can be recovered;
|
||||
- how every supported update path is tested.
|
||||
|
||||
A package manager alone solves only part of the OTA problem.
|
||||
|
||||
---
|
||||
|
||||
## 5. Container-Based Updates
|
||||
|
||||
A container image may be used as the delivery and execution unit for an application.
|
||||
|
||||
In this model, the base operating system provides a container runtime, while application components are delivered separately.
|
||||
|
||||
```text
|
||||
Host operating system
|
||||
├── Container A
|
||||
├── Container B
|
||||
└── Container C
|
||||
```
|
||||
|
||||
Updating an application may consist of downloading a new image and switching to it.
|
||||
|
||||
### 5.1. Advantages
|
||||
|
||||
- the application is delivered with a significant part of its dependencies;
|
||||
- strong component isolation;
|
||||
- convenient versioning;
|
||||
- relatively simple return to a previous image;
|
||||
- reuse of existing build and registry infrastructure;
|
||||
- applications can have independent release cycles.
|
||||
|
||||
### 5.2. Disadvantages
|
||||
|
||||
- the container runtime becomes part of the trusted platform;
|
||||
- storage and memory requirements increase;
|
||||
- another operational layer is introduced;
|
||||
- application state still requires separate management;
|
||||
- container updates do not update the kernel, drivers, or base operating system;
|
||||
- desktop or cloud practices cannot automatically be transferred to a constrained embedded device.
|
||||
|
||||
Containers do not replace OTA architecture. They may become one layer of it.
|
||||
|
||||
For example:
|
||||
|
||||
```text
|
||||
Base OS -> A/B system images
|
||||
Applications -> containers
|
||||
Configuration -> signed documents
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. Atomic Filesystem Trees and Versioned System States
|
||||
|
||||
Between full images and traditional packages are approaches in which the system is represented as a versioned filesystem tree.
|
||||
|
||||
A new version is assembled separately, after which the device atomically switches to it. Unchanged objects may be reused, so a fully independent image does not always need to be transferred or stored.
|
||||
|
||||
Concepts in this class include:
|
||||
|
||||
- content-addressed storage;
|
||||
- immutable filesystem trees;
|
||||
- snapshot-based deployments;
|
||||
- OSTree-like models.
|
||||
|
||||
The central idea is:
|
||||
|
||||
> The update is neither an in-place set of file changes nor necessarily a complete block image. It is a complete versioned filesystem state.
|
||||
|
||||
### 6.1. Advantages
|
||||
|
||||
- atomic switching between versions;
|
||||
- reproducible system state;
|
||||
- reuse of unchanged content;
|
||||
- convenient rollback to a previous tree;
|
||||
- fewer changes to the running system;
|
||||
- a useful compromise between images and packages.
|
||||
|
||||
### 6.2. Disadvantages
|
||||
|
||||
- more complex storage model;
|
||||
- additional build-tooling requirements;
|
||||
- garbage collection and multi-version management;
|
||||
- state outside the managed tree requires separate handling;
|
||||
- integration with the bootloader and early boot may be non-trivial;
|
||||
- the approach may be excessive for a simple device.
|
||||
|
||||
This mechanism can be attractive, but its benefits appear only when the entire system lifecycle is genuinely designed around immutable, versioned states.
|
||||
|
||||
---
|
||||
|
||||
## 7. Binary Delta Updates
|
||||
|
||||
A binary delta contains the difference between a known old version and a new version rather than the complete new artifact.
|
||||
|
||||
Conceptually:
|
||||
|
||||
```text
|
||||
Old artifact + Delta = New artifact
|
||||
```
|
||||
|
||||
This may significantly reduce the amount of transferred data, especially when only a small portion of the artifact has changed.
|
||||
|
||||
However, a delta is not a separate answer to the question of what is being updated.
|
||||
|
||||
It may be applied to:
|
||||
|
||||
- a system image;
|
||||
- a partition;
|
||||
- a package;
|
||||
- a container layer;
|
||||
- an individual file;
|
||||
- a model or data set.
|
||||
|
||||
It is therefore more accurate to treat a binary delta as a **delivery optimization**, not as an update unit.
|
||||
|
||||
### 7.1. Advantages
|
||||
|
||||
- lower network traffic;
|
||||
- faster download over slow connections;
|
||||
- reduced mobile or satellite data cost;
|
||||
- the ability to update large artifacts through a constrained channel.
|
||||
|
||||
### 7.2. Disadvantages
|
||||
|
||||
- the delta depends on a specific source version;
|
||||
- multiple update paths may need to be stored or generated;
|
||||
- corruption of the source artifact may make application impossible;
|
||||
- the device requires resources to reconstruct the new version;
|
||||
- both server-side and device-side complexity increase;
|
||||
- the delta must be verified as strictly as the full artifact;
|
||||
- savings may be small when changes are poorly localized.
|
||||
|
||||
For example, updating from version `1.0` to `1.1` may require one delta, while updating from `0.8` to `1.1` requires another.
|
||||
|
||||
With many supported versions, update paths multiply:
|
||||
|
||||
```text
|
||||
0.8 -> 1.1
|
||||
0.9 -> 1.1
|
||||
1.0 -> 1.1
|
||||
1.0 -> 1.0.1 -> 1.1
|
||||
```
|
||||
|
||||
The more possible paths there are, the harder they are to create, test, and maintain.
|
||||
|
||||
Reduced network traffic is therefore purchased with additional system complexity.
|
||||
|
||||
---
|
||||
|
||||
## 8. Updating Individual Files and Data
|
||||
|
||||
The smallest update unit is an individual file or a small group of files.
|
||||
|
||||
This is a natural approach for:
|
||||
|
||||
- configuration;
|
||||
- certificates;
|
||||
- rules;
|
||||
- calibration data;
|
||||
- static content;
|
||||
- models;
|
||||
- maps;
|
||||
- user-interface resources.
|
||||
|
||||
At first glance, such an update seems simple: download a file and replace the old one.
|
||||
|
||||
A reliable implementation still needs to answer:
|
||||
|
||||
- how the version is checked;
|
||||
- how authenticity is verified;
|
||||
- how integrity is verified;
|
||||
- how compatibility is verified;
|
||||
- how a partially written file is avoided;
|
||||
- how replacement is made atomic;
|
||||
- what happens if power is lost;
|
||||
- how the previous version is preserved;
|
||||
- who confirms successful activation.
|
||||
|
||||
Even a small file can render the device unusable if it contains critical configuration.
|
||||
|
||||
Artifact size does not determine the cost of failure.
|
||||
|
||||
---
|
||||
|
||||
## 9. Hybrid Strategies
|
||||
|
||||
In most complex products, the most reasonable solution is a combination of several mechanisms.
|
||||
|
||||
For example:
|
||||
|
||||
```text
|
||||
Bootloader and base OS
|
||||
-> signed full images
|
||||
-> A/B installation
|
||||
-> reboot-based activation
|
||||
|
||||
Applications
|
||||
-> packages or containers
|
||||
-> independent release cycle
|
||||
|
||||
Maps and ML models
|
||||
-> separate artifacts
|
||||
-> optional delta delivery
|
||||
|
||||
Configuration
|
||||
-> small signed documents
|
||||
-> atomic replacement
|
||||
|
||||
Certificates
|
||||
-> dedicated rotation protocol
|
||||
```
|
||||
|
||||
This allows different properties to be applied to different components.
|
||||
|
||||
For the base system, the priority may be reliability and reproducibility. For maps, it may be minimal network traffic. For applications, rapid release. For certificates, strict coordination and limited validity periods.
|
||||
|
||||
### 9.1. Advantages
|
||||
|
||||
- each data type receives an appropriate mechanism;
|
||||
- the entire system does not need to be updated for a small change;
|
||||
- reliability, bandwidth, and release speed can be optimized independently;
|
||||
- the design reflects the real lifecycle of each component.
|
||||
|
||||
### 9.2. Disadvantages
|
||||
|
||||
- multiple mechanisms must be designed, tested, and maintained;
|
||||
- compatibility modelling becomes more complex;
|
||||
- identifying the complete product version becomes more difficult;
|
||||
- dependencies may exist between separate update channels;
|
||||
- rolling back one component may require rolling back others;
|
||||
- observability and diagnostics become more complex.
|
||||
|
||||
A hybrid strategy must not become an accidental collection of unrelated mechanisms.
|
||||
|
||||
Each independent update channel should exist for a concrete reason:
|
||||
|
||||
- a different lifecycle;
|
||||
- a different transfer cost;
|
||||
- a different cost of failure;
|
||||
- different activation requirements;
|
||||
- a need for organizational independence.
|
||||
|
||||
Without such a reason, an additional mechanism only increases system complexity.
|
||||
|
||||
---
|
||||
|
||||
## 10. Why There Is No Best Strategy
|
||||
|
||||
The question:
|
||||
|
||||
> Which OTA strategy is best?
|
||||
|
||||
has no universal answer.
|
||||
|
||||
A full image may be an excellent choice for an industrial controller and too expensive for a device using satellite connectivity.
|
||||
|
||||
Packages may fit naturally into a Linux gateway and create an unacceptable number of states in a safety-critical system.
|
||||
|
||||
Containers may simplify independent application releases while adding a pointless operational layer to a small device.
|
||||
|
||||
A binary delta may dramatically reduce transfer cost while introducing more complexity than the rest of the updater.
|
||||
|
||||
The correct question is:
|
||||
|
||||
> **What constraints does our system have, and which update properties matter most?**
|
||||
|
||||
### 10.1. Reliability
|
||||
|
||||
What happens if power is lost in the middle of installation?
|
||||
|
||||
Must the device always return to the previous version?
|
||||
|
||||
Is manual recovery acceptable?
|
||||
|
||||
The more expensive physical access is, the more valuable atomic installation and automatic rollback become.
|
||||
|
||||
### 10.2. Storage Capacity
|
||||
|
||||
Is there enough space for two complete system images?
|
||||
|
||||
Can the device store the downloaded artifact while also constructing the new version?
|
||||
|
||||
How many previous versions must be retained?
|
||||
|
||||
Limited storage may exclude some approaches or require a more complex streaming installation process.
|
||||
|
||||
### 10.3. Connectivity Cost and Quality
|
||||
|
||||
Does the device use Ethernet, Wi-Fi, mobile, or satellite connectivity?
|
||||
|
||||
Is the connection stable or frequently interrupted?
|
||||
|
||||
Is traffic expensive?
|
||||
|
||||
Can downloads safely resume?
|
||||
|
||||
Over an expensive link, the complexity of delta updates may be justified. On a cheap local network, a full image may be simpler and cheaper to operate.
|
||||
|
||||
### 10.4. Update Frequency
|
||||
|
||||
Is the system updated once per year or several times per day?
|
||||
|
||||
Rare system updates and frequent data updates naturally require different mechanisms.
|
||||
|
||||
The more frequently a component changes, the more valuable an independent release cycle becomes.
|
||||
|
||||
### 10.5. Component Independence
|
||||
|
||||
Can the components really be updated independently?
|
||||
|
||||
Are their interfaces stable?
|
||||
|
||||
Is compatibility between versions maintained?
|
||||
|
||||
Independent delivery does not imply independent operation. If application version `5` requires library version `3`, the two must still be coordinated.
|
||||
|
||||
### 10.6. Cost of Failure
|
||||
|
||||
What happens if an update is unusable?
|
||||
|
||||
Does the device temporarily lose a minor feature? Stop a production line? Make a vehicle unavailable? Require a technician visit?
|
||||
|
||||
The same technical failure may have completely different consequences in different products.
|
||||
|
||||
### 10.7. Downtime Requirements
|
||||
|
||||
Can the device be rebooted?
|
||||
|
||||
How long may version activation take?
|
||||
|
||||
Is restarting one service acceptable?
|
||||
|
||||
Some systems may update overnight with a full reboot. Others must preserve their primary function almost continuously.
|
||||
|
||||
### 10.8. Support Lifetime
|
||||
|
||||
How many years will the device remain in service?
|
||||
|
||||
Must the system support updates from very old versions?
|
||||
|
||||
Will a team still be available to maintain a complex mechanism ten years from now?
|
||||
|
||||
The OTA architecture must not only be technically possible today. It must remain maintainable throughout the entire product lifecycle.
|
||||
|
||||
### 10.9. Team Capabilities
|
||||
|
||||
A complex solution does not become a good solution merely because it is technically elegant.
|
||||
|
||||
The team must be able to:
|
||||
|
||||
- build artifacts;
|
||||
- test update paths;
|
||||
- operate server infrastructure;
|
||||
- investigate failed updates;
|
||||
- manage cryptographic keys;
|
||||
- maintain compatibility;
|
||||
- recover devices.
|
||||
|
||||
If potential traffic savings require a mechanism the team cannot operate reliably, the optimization may increase the total cost of the product.
|
||||
|
||||
---
|
||||
|
||||
## 11. Strategy Is a Consequence of Constraints
|
||||
|
||||
OTA selection should not begin with a technology name.
|
||||
|
||||
It should begin with a model of the product:
|
||||
|
||||
```text
|
||||
What is updated?
|
||||
|
|
||||
v
|
||||
How often is it updated?
|
||||
|
|
||||
v
|
||||
How expensive is failure?
|
||||
|
|
||||
v
|
||||
What resources are available?
|
||||
|
|
||||
v
|
||||
What recovery guarantees are required?
|
||||
|
|
||||
v
|
||||
Which mechanism satisfies these constraints?
|
||||
```
|
||||
|
||||
Therefore:
|
||||
|
||||
> **An OTA strategy is a consequence of system constraints, not a choice of the newest or most popular technology.**
|
||||
|
||||
The engineering process should look approximately like this:
|
||||
|
||||
1. Identify every independently updateable component.
|
||||
2. Describe the lifecycle of each component.
|
||||
3. Determine the cost of a failed update.
|
||||
4. Record storage, memory, power, and connectivity constraints.
|
||||
5. Define atomicity, activation, and rollback requirements.
|
||||
6. Select an update unit for each component.
|
||||
7. Only then choose specific technologies and tools.
|
||||
|
||||
The order matters.
|
||||
|
||||
If the process starts with a tool, the architecture will be forced to fit the tool's capabilities. If it starts with constraints, the tool can be selected for the actual problem.
|
||||
|
||||
---
|
||||
|
||||
## 12. Preliminary Comparison
|
||||
|
||||
The following table is not yet a complete decision matrix. It only summarizes the primary properties of the approaches discussed in this chapter.
|
||||
|
||||
| Approach | Primary Unit | Main Strength | Main Cost |
|
||||
|---|---|---|---|
|
||||
| Full image | System or partition | Reproducibility and predictable rollback | Artifact size and storage requirements |
|
||||
| Packages | System component | Independent component updates | Many possible states and difficult rollback |
|
||||
| Application | Individual service or bundle | Fast independent releases | Platform compatibility |
|
||||
| Container | Application with dependencies | Isolation and controlled delivery | Additional runtime and operational complexity |
|
||||
| Filesystem tree | Complete filesystem state | Atomicity with content reuse | More complex storage model |
|
||||
| Binary delta | Difference between artifact versions | Reduced network traffic | Dependence on the source version |
|
||||
| Individual files | Configuration or data | Small and targeted changes | Atomicity and compatibility must be designed explicitly |
|
||||
| Hybrid model | Multiple update units | Optimization for different lifecycles | Overall system complexity |
|
||||
|
||||
There is intentionally no *best option* column.
|
||||
|
||||
The same approach may be strong or weak depending on context.
|
||||
|
||||
---
|
||||
|
||||
## 13. What Comes Next
|
||||
|
||||
This chapter describes the space of possible solutions, but it does not declare a single winner.
|
||||
|
||||
The next step will be a series of Architecture Decision Records in which each option is evaluated in the context of specific constraints.
|
||||
|
||||
Each ADR should answer not:
|
||||
|
||||
> What is this technology?
|
||||
|
||||
but:
|
||||
|
||||
> **Under which conditions would I choose it, and which consequences of that decision would I be prepared to accept?**
|
||||
|
||||
The following topics will be considered separately:
|
||||
|
||||
- when to choose full-image updates;
|
||||
- when to choose package-based updates;
|
||||
- when to update individual applications;
|
||||
- when to use containers;
|
||||
- when binary deltas are justified;
|
||||
- when versioned filesystem trees are appropriate;
|
||||
- how to design a hybrid strategy.
|
||||
|
||||
The results will then be summarized in a decision matrix.
|
||||
|
||||
The matrix should not select a solution automatically. Its purpose is to show which constraints push the architecture toward a particular approach and where trade-offs begin to conflict.
|
||||
|
||||
---
|
||||
|
||||
## Conclusion
|
||||
|
||||
An OTA update cannot be reduced to a choice between full images and packages.
|
||||
|
||||
A real system must separately define:
|
||||
|
||||
- what is updated;
|
||||
- how independently it is updated;
|
||||
- how it is delivered;
|
||||
- where it is prepared;
|
||||
- how it is activated;
|
||||
- how it is verified;
|
||||
- how it is rolled back;
|
||||
- what happens when something fails.
|
||||
|
||||
Full images, packages, containers, filesystem trees, and individual files define different units of system management. A/B layouts, deltas, signatures, integrity checks, and recovery mechanisms add the properties required around those units.
|
||||
|
||||
A good OTA design therefore does not begin with a technology.
|
||||
|
||||
It begins with constraints, the product lifecycle, and an honest answer to the question:
|
||||
|
||||
> **What device state must we guarantee after every possible update outcome?**
|
||||
27
readme.md
27
readme.md
@@ -1,5 +1,28 @@
|
||||
# ota-reference-design
|
||||
|
||||
A practical reference design for reliable and secure over-the-air updates on embedded Linux
|
||||
A practical reference design for reliable and secure over-the-air updates on embedded Linux.
|
||||
|
||||
This repository is intended as an educational reference design. It demonstrates one possible approach to building a reliable OTA update system for embedded Linux. It is not intended to be a production-ready framework.
|
||||
This repository is intended as an educational reference design. It demonstrates one possible approach to designing and implementing a reliable OTA update system for embedded Linux.
|
||||
|
||||
The project focuses on the engineering decisions behind OTA systems, including:
|
||||
|
||||
* update strategies;
|
||||
* system and storage architecture;
|
||||
* bootloader interaction;
|
||||
* atomic updates;
|
||||
* rollback and recovery;
|
||||
* integrity verification and signing;
|
||||
* testing and failure handling.
|
||||
|
||||
It is not tied to a specific OTA framework, hardware platform, or cloud provider.
|
||||
|
||||
Raspberry Pi may be used for practical demonstrations, but the underlying concepts are applicable to a broader range of embedded Linux devices.
|
||||
|
||||
This repository is not intended to be a production-ready framework. Instead, it is designed to explain the architectural principles, trade-offs, and failure scenarios that should be considered when building a real OTA system.
|
||||
|
||||
The repository is being developed incrementally as a structured series of articles, diagrams, architecture decisions, and practical examples.
|
||||
|
||||
|
||||
## License
|
||||
|
||||
This project is licensed under the MIT License. See [LICENSE](LICENSE) for details.
|
||||
|
||||
Reference in New Issue
Block a user