Object storage: the definition
Object storage manages data as individually addressed objects, usually combining content with metadata and a key inside a container such as a bucket. Access is generally through a service API rather than direct block reads and writes.
The key points
- An object key identifies data; a folder-like prefix is not necessarily a filesystem directory.
- Consistency and durability guarantees must be read for the specific service.
- Versioning is useful but does not replace permission design and tested recovery.
- Request size, concurrency, caching and transfer charges can matter as much as stored capacity.
What is stored as an object?
An object combines a sequence of bytes with an identifier and associated metadata. In Amazon S3, for example, a bucket and key identify an object, with a version identifier distinguishing retained versions where versioning is used. The key can look like a path, but the storage interface is an object service. This distinction affects how applications list, replace and retrieve data. [1]
Imagine a hypothetical training dataset containing an object named images/train/shard-001.bin. The slashes make its key convenient to organize, but do not by themselves prove that three ordinary filesystem directories exist underneath it. Treating every object store as a mounted local filesystem can hide differences in rename behavior, access granularity and consistency. Use the service's documented contract rather than the visual appearance of its browser.
Object, file and block storage
Block storage exposes addressable blocks that a higher layer can organize. File storage exposes files and filesystem operations. Object storage exposes objects through a service API. These models overlap in what data they can ultimately hold, but differ in access semantics. A compatibility layer can make one resemble another without making every operation equivalent in cost, latency or behavior. [2]
For an illustrative choice, a database expecting frequent in-place page updates and filesystem guarantees should not be moved to an arbitrary object interface simply because both can hold bytes. A collection of large immutable dataset shards may be a much more natural fit. Choose the access model around the software's behavior, not a rule that one storage category is universally superior.
Keys, versions and reproducible datasets
A human-readable key can refer to different contents over time if objects are replaced. Versioning can retain distinct versions, and content hashes can provide another way to identify exact bytes. Reproducibility requires recording which data a job actually consumed, not just the latest key name. The same principle applies to model checkpoints and software images. [1][3]
Suppose a hypothetical experiment records only bucket/dataset/latest.bin. Another team replaces that object before the experiment is repeated. The path is unchanged, but the input is not. Recording a version or immutable manifest makes the experiment much more interpretable. A manifest should also identify preprocessing and sampling assumptions; immutable raw bytes alone do not define every aspect of a training run.
We recommend treating dataset manifests as first-class experiment inputs. A manifest can list exact object identities, expected sizes and integrity checks. It then becomes possible to distinguish a changed model implementation from changed data. This is useful even when the storage service already provides strong consistency, because consistency and historical identity solve different problems.
Consistency is a service contract
Consistency describes what clients can observe after operations. Amazon S3 documents strong read-after-write consistency for specified object operations, including replacements and deletions, and atomic updates to an individual key. That is a specific service guarantee. It should not be generalized to every product with an S3-compatible API or to an arbitrary multi-object transaction. [2]
Imagine a hypothetical dataset update involving a thousand objects and a manifest. Even if each single-object write is atomic, readers need a publication protocol that prevents them from treating a half-uploaded collection as complete. One approach is to write immutable objects first and publish the manifest only after verification. The application must still define how readers select and validate that manifest.
Throughput, requests and concurrency
Object-service performance depends on how applications issue requests. Concurrency, object sizing and the client environment influence throughput; provider guidance should be read for the chosen service. A large sequential object transfer differs from millions of tiny requests. The bottleneck can be request overhead, CPU processing, client networking, service limits or downstream consumption. [4]
Suppose a hypothetical application makes 10,000 serial requests, each taking 20 milliseconds before useful processing. That path alone consumes about 200 seconds. Packing data into larger shards or safely overlapping requests may help, but can change random-access efficiency and retry granularity. Bigger objects are not always better; the layout should reflect how the application reads and recovers.
A cache can also change the experiment. A second training epoch reading from a local cache may no longer measure the remote object's storage path. Record cold-start behavior separately from steady-state processing. Otherwise a benchmark can appear to prove remote throughput while actually measuring local memory or SSD performance.
Permissions, versioning and recovery
Object stores expose access controls at service-defined boundaries. In S3, policies and identity permissions determine access, and the documentation recommends modern policy-based arrangements rather than assuming every bucket needs legacy ACLs. A public-readable URL is a security decision, not a necessary feature of object storage. Credentials should grant only the operations and data scope required by the workload. [5]
Versioning can help preserve earlier states after replacement or deletion, but its behavior must be understood alongside lifecycle policies, retention and administrative permissions. A powerful actor or an unsuitable policy can still undermine recovery. Test a restore using the same access context and procedure that would be available during an incident. [3][6]
As a hypothetical failure drill, assume a job accidentally overwrites a checkpoint key. Determine whether an earlier version remains, who may retrieve it and how the application selects it. Then consider a separate credential compromise or bucket-deletion scenario. A successful answer to the first case does not automatically cover the others.
Stored capacity is not the whole bill
A storage-cost model should identify the applicable service's capacity, request, retrieval and data-transfer terms, together with any minimum duration or class-specific conditions. This chapter does not supply a current universal tariff. The important method is to count the operations and movement created by the workload, not assume that a per-GB-month figure describes the entire cost.
If a hypothetical service stores 5 TB but reads all of it daily into another environment, monthly data movement can be much larger than stored capacity. A cheap storage line item can therefore coexist with an expensive pipeline. Conversely, a reusable local cache may reduce transfers while adding local storage cost and invalidation responsibilities. Model the complete path using actual terms.
A clear architecture decision
Object storage is especially understandable when data identity, access patterns and publication rules are explicit. Decide which objects are immutable, how readers find a complete version, how integrity is checked, what permissions are needed and how recovery is tested. Then measure the actual ingestion and checkpoint workload under realistic concurrency.
The core lesson is that an object store is a service with data and operation contracts. It can be a strong foundation for datasets and model artifacts, but it is not automatically a local disk, a transactional filesystem or a complete backup strategy. Keeping those distinctions visible prevents both performance surprises and recovery mistakes.
Check your understanding
Try answering before opening the explanation. Your answers are not collected or scored.
1. Does a slash in an object key guarantee a real directory?
No. It can be part of the key or a listing prefix. Filesystem semantics must be checked rather than inferred from the name.
2. Does atomic replacement of one object guarantee an atomic dataset update?
No. A dataset can span many objects and needs an application-level publication protocol, such as verified immutable objects followed by a manifest.
3. Why record versions or content identities?
A key can be reused for different bytes. Exact identities help reproduce experiments and distinguish data changes from software changes.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- AWS · Amazon S3 objects overview ↗ (opens in a new tab)Provider documentation · Checked 28 September 2026
- AWS · S3 architecture and consistency ↗ (opens in a new tab)Provider documentation · Checked 28 September 2026
- AWS · S3 Versioning ↗ (opens in a new tab)Provider documentation · Checked 28 September 2026
- AWS · S3 performance design patterns ↗ (opens in a new tab)Provider documentation · Checked 28 September 2026
- AWS · Object access controls ↗ (opens in a new tab)Provider documentation · Checked 28 September 2026
- AWS · Managing object lifecycle ↗ (opens in a new tab)Provider documentation · Checked 28 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.