---
title: "Object Storage Index Output"
description: "index files uploaded to a object storage container (e.g. AWS S3 bucket)."
source: "https://github.com/log-10x/modules/tree/main/pipelines/run/modules/input/objectStorage/index/module.yaml"
icon: "material/target"

---
The Index module enables [storage queries](https://doc.log10x.com/run/input/objectStorage/query "Query an object storage container for events matching target criteria") to fetch blob byte ranges (e.g., [AWS S3](https://docs.aws.amazon.com/whitepapers/latest/s3-optimizing-performance-best-practices/use-byte-range-fetches.html){target="\_blank"}, [Azure](https://learn.microsoft.com/en-us/rest/api/storageservices/specifying-the-range-header-for-blob-service-operations){target="\_blank"})
matching a specified app/service name, search term(s) and timestamp range predictably at scale.

### :material-hammer-wrench: Workflow

The index [module](https://doc.log10x.com/engine/module/) executes when S3 event notifications are sent directly to [SQS queues](https://doc.log10x.com/apps/retriever/#sqs-based-architecture), triggering index workers to process uploaded files.

The module comprises an _input_ and _output_ stream:

- The [input stream](https://github.com/log-10x/modules/blob/main/pipelines/run/modules/input/objectStorage/query/stream.yaml){target="\_blank"} reads the events from the uploaded file to transform them into [TenXObjects](https://doc.log10x.com/api/js/#TenXObject "Provide structured, reflective access to log/trace events read from input(s).").

- The [output stream](https://github.com/log-10x/modules/blob/main/pipelines/run/modules/input/objectStorage/index/stream.yaml){target="\_blank"} performs the following actions for each TenXObject:

1. Write its [template](https://doc.log10x.com/api/js/#TenXBaseObject+template "A sequence of all symbol and delimiter tokens from the object's text field.") to the [index](#indexwritecontainer "name of target index container") container (if not exists).
2. Map its [timestamp](https://doc.log10x.com/api/js/#TenXObject+timestamp "An array of UNIX epoch values of timestamps parsed from the object's text.") to a [Bloom filter](https://en.wikipedia.org/wiki/Bloom_filter){target="\_blank"} associated with a rolling time window specified by [indexWriteResolution](#indexwriteresolution "index time window resolution").
3. Append its [templateHash](https://doc.log10x.com/api/js/#TenXBaseObject+templateHash "An alphanumeric encoded value of a 64bit hash of the current object's template field.") and [vars](https://doc.log10x.com/api/js/#TenXBaseObject+vars "An array of  variable sequences extracted from the object's text.") to the Bloom filter's hash set. Once a filter's size exceeds the Object storage's key byte length (e.g., for AWS S3 [1024 bytes](https://docs.aws.amazon.com/AmazonS3/latest/userguide/object-keys.html){target="\_blank"}), the stream writes it to the [index](#indexwritecontainer "name of target index container") container and assigns a new filter to the time window.

### :material-filter-variant: TenXTemplate Filters

TenXTemplate Bloom filters enable parallel traversal of the [index container](#indexwritecontainer "name of target index container") (e.g., S3 bucket) with high [test accuracy](https://doc.log10x.com/run/input/objectStorage/index/#accuracy) for fetching required byte ranges to scan for matching log/trace events.

Separating low-cardinality [symbol](https://doc.log10x.com/run/transform/structure/#symbols) values into TenXTemplates and writing only template hashes and high-cardinality [variables](https://doc.log10x.com/run/transform/structure/#variables) to Bloom filters reduces their volume by over 75% compared to appending both low and high-cardinality values.

Restricting Bloom filter size to the object storage's key length enables batch retrieval of filters via *list* operations (e.g., [AWS S3](https://docs.aws.amazon.com/AmazonS3/latest/API/API_ListObjects.html){target="\_blank"}: 1000 keys/request, [Azure](https://learn.microsoft.com/en-us/rest/api/storageservices/list-blobs?tabs=microsoft-entra-id){target="\_blank"}: 5000 keys/request, [GCP](https://cloud.google.com/storage/docs/json_api/v1/objects/list){target="\_blank"}: 1000 keys/request).

```mermaid
graph LR
    A["📥 Log File Upload<br/>(S3/Azure/GCP)"] --> B["Read Events<br/>Stream"]
    B --> C["Transform to<br/>TenXObjects"]
    C --> D["⚙️ Process<br/>Each Object"]

    D --> E["Extract Variables<br/>& Template"]
    E --> F["Map to<br/>Time Window"]
    F --> G{"Filter Size<br/>< 1024 bytes?"}
    
    G -->|Yes| H["Append to<br/>Current Filter"]
    G -->|No| I["Write Filter<br/>to Index"]
    
    H --> J{"More<br/>Events?"}
    I --> K["Create New<br/>Filter"]
    K --> J
    
    J -->|Yes| D
    J -->|No| L["✅ Index Objects<br/>Ready for Query"]
    
    E --> M["Write Template<br/>(if new)"]
    M --> L
    
    classDef input fill:#2563eb88,stroke:#1d4ed8,color:#ffffff,stroke-width:2px,rx:8,ry:8
    classDef processing fill:#05966988,stroke:#047857,color:#ffffff,stroke-width:2px,rx:8,ry:8
    classDef output fill:#dc262688,stroke:#b91c1c,color:#ffffff,stroke-width:2px,rx:8,ry:8
    classDef decision fill:#7c3aed88,stroke:#6d28d9,color:#ffffff,stroke-width:2px,rx:8,ry:8
    classDef result fill:#ea580c88,stroke:#c2410c,color:#ffffff,stroke-width:2px,rx:8,ry:8
    
    class A,B,C input
    class D,E,F processing
    class G,J decision
    class H,I,K,M output
    class L result

```

<div class="diagram-controls">
    <button class="md-button md-button--primary enlarge-diagram" onclick="enlargeDiagram(this)" data-tooltip="Click to enlarge diagram">
        <span class="twemoji">
            <svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24" width="16" height="16">
                <path d="M10 2c4.42 0 8 3.58 8 8 0 1.85-.63 3.55-1.69 4.9L20.59 19l-1.41 1.41-4.09-4.09A7.84 7.84 0 0 1 10 18c-4.42 0-8-3.58-8-8s3.58-8 8-8m0 2a6 6 0 1 0 0 12 6 6 0 0 0 0-12m1 3h2v2h-2V7m-4 0h2v2H7V7m2 4h2v2H9v-2Z"/>
            </svg>
        </span>
        Enlarge Diagram
    </button>
</div>
<!-- Mermaid enhanced diagram functionality loaded via external files -->

### :material-server-outline: Compute Resources

Indexing is CPU and memory intensive during file parsing. Default k8s pod resources:

- **1 CPU** and **2GB memory** per pod (see [deployment guide](https://doc.log10x.com/apps/retriever/deploy/#step-4-configure-application))
- **Autoscaling:** 2–10 replicas depending on queue depth (default 2 min, scales to 10 if backlog grows)
- **Throughput:** One pod handles ~10–50 GB/day depending on event size and CPU availability

Indexing runs asynchronously, triggered by S3 event notifications, in parallel with queries. Multiple index workers process files concurrently from the SQS queue. Indexes are built once at ingest time and never recomputed.

### :material-currency-usd: Cost

Index building cost is part of the k8s pod resource costs, no per-GB indexing fee. You pay:

- k8s pod (CPU + memory) running the index workers
- S3 storage for index objects (~1–5% overhead vs. original data size)
- SQS queue operations (~$0.40 per million messages)

### :material-arrow-up-bold: Scaling

If files upload faster than indexing, the SQS queue buffers pending work, no events are lost. Index worker pods scale up automatically via Kubernetes HPA.

Unindexed files remain queryable via full scan (slower than indexed queries but functional).

**Deployment topologies:**

- **All-in-one:** Single pod cluster handles index, query, and stream roles (suitable for \<100 GB/day)
- **Separate clusters:** Dedicated index/query/stream pods allow independent scaling (recommended for \>500 GB/day)

See the [deployment guide](https://doc.log10x.com/apps/retriever/deploy/#step-4-configure-application) for sizing guidance.

## :material-wrench-outline: Config Files

To configure the Object storage index output module, [:material-cog: Edit](https://doc.log10x.com/config/app/#module-config "Learn how to edit app and module configurations") these files.  

## :material-menu: Options

Specify the options below to [configure](/config "configure") multiple Object storage index output:

|Name|Description|Category|
|---|---|---|
|[indexObjectStorageName](#indexobjectstoragename "object storage logical name")|Object storage logical name|Container|
|[indexReadContainer](#indexreadcontainer "name of the input Object Storage container")|Name of the input Object Storage container|Container|
|[indexReadObject](#indexreadobject "name of the object storage blob to index")|Name of the object storage blob to index|Container|
|[indexWriteContainer](#indexwritecontainer "name of target index container")|Name of target index container|Container|
|[indexWriteTarget](#indexwritetarget "logical name identifying the origin of 'indexReadObject'")|Logical name identifying the origin of 'indexReadObject'|Output|
|[indexReadExtractMessage](#indexreadextractmessage "use extractor for inner message")|Use extractor for inner message|Parsing|
|[indexReadMessageField](#indexreadmessagefield "message field name")|Message field name|Parsing|
|[indexWriteByteRange](#indexwritebyterange "max byte range size to index 'indexReadObject'")|Max byte range size to index 'indexReadObject'|Accuracy|
|[indexWriteResolution](#indexwriteresolution "index time window resolution")|Index time window resolution|Accuracy|
|[indexWriteAccuracy](#indexwriteaccuracy "bloom filter accuracy of index objects.")|Bloom filter accuracy of index objects.|Accuracy|
|[indexWriteTemplateMergeInterval](#indexwritetemplatemergeinterval "merge template interval")|Merge template interval|Advanced|
|[indexObjectStorageArgs](#indexobjectstorageargs "custom Object storage args")|Custom Object storage args|Advanced|
|[indexReadPrintProgress](#indexreadprintprogress "sets whether this input prints throughput stats to the console")|Sets whether this input prints throughput stats to the console|General|

### Container

#### :material-menu-right-outline:**`indexObjectStorageName`**

Object storage logical name.

|Type|Required|Category|
|---|---|---|
|String|✔|Container|

Identifies the [Object Storage](https://doc.log10x.com/run/input/objectStorage#/objectstorageaccessclassname) containing the blob to index (e.g., AWS).


#### :material-menu-right-outline:**`indexReadContainer`**

Name of the input Object Storage container.

|Type|Required|Category|
|---|---|---|
|String|✔|Container|

Specifies the Object Storage container (e.g., AWS S3 bucket) containing the blob (e.g., log file) to index.


#### :material-menu-right-outline:**`indexReadObject`**

Name of the object storage blob to index.

|Type|Required|Category|
|---|---|---|
|String|✔|Container|

Specifies the blob (e.g., log file) name within [indexReadContainer](https://doc.log10x.com/run/input/objectStorage/index/#indexreadcontainer "name of the input Object Storage container") to index.


#### :material-menu-right-outline:**`indexWriteContainer`**

Name of target index container.

|Type|Required|Category|
|---|---|---|
|String|✔|Container|

Specifies the storage container (e.g., AWS S3 bucket) to output [TenXTemplate Filters](https://doc.log10x.com/run/input/objectStorage/index/#tenxtemplate-filters)  (e.g., TenXTemplates and Bloom filters).


### Output

#### :material-menu-right-outline:**`indexWriteTarget`**

Logical name identifying the origin of 'indexReadObject'.

|Type|Required|Category|
|---|---|---|
|String|✔|Output|

Specifies a logical name to store index objects produced for [indexReadObject](https://doc.log10x.com/run/input/objectStorage/index/#indexreadobject "name of the object storage blob to index") under.
This name commonly specifies the app which generated the events enclosed within this blob (e.g. `acme-client`).


### Parsing

#### :material-menu-right-outline:**`indexReadExtractMessage`**

Use extractor for inner message.

|Type|Default|Category|
|---|---|---|
|Boolean|false|Parsing|

Specifies whether to extract an inner field from the entire event json to use as the base for constructing the TenXObject.


#### :material-menu-right-outline:**`indexReadMessageField`**

Message field name.

|Type|Default|Category|
|---|---|---|
|String|log|Parsing|

Name of the actual message field in the event json to use to construct the TenXObject, used only if [indexReadExtractMessage](https://doc.log10x.com/run/input/objectStorage/index/#indexreadextractmessage "use extractor for inner message") is true.


### Accuracy

#### :material-menu-right-outline:**`indexWriteByteRange`**

Max byte range size to index 'indexReadObject'.

|Type|Required|Category|
|---|---|---|
|Number|✔|Accuracy|

Controls the chunk size in which to index the target object.

For example, if the target object is 1GB and this value is `2MB`,
index the object in 2MB segments to ensure matching queries
can retrieve chunks vs. all of it unnecessarily.
To learn more see: [byte range fetches](https://docs.aws.amazon.com/whitepapers/latest/s3-optimizing-performance-best-practices/use-byte-range-fetches.html){target="\_blank"}.


#### :material-menu-right-outline:**`indexWriteResolution`**

Index time window resolution.

|Type|Required|Category|
|---|---|---|
|Number|✔|Accuracy|

Controls the index time range resolution.

For example, setting this to `1min` means that queries to the index at time
ranges greater than 1min (e.g. 15min) will not fetch byte ranges
outside the time frame unnecessarily.

The lower this value is, the greater the output index size will be.

This value should satisfy the minimum resolution for querying the index.
For example, if queries to the index are in 5-minute increments:

```yaml
query:
  filter:
    from: $=now("-5m")
    to: $=now()
```

Setting this value to `5min` will create the most efficient index.


#### :material-menu-right-outline:**`indexWriteAccuracy`**

Bloom filter accuracy of index objects.

|Type|Required|Category|
|---|---|---|
|Number|✔|Accuracy|

Controls the accuracy of bloom filter [TenXTemplate Filters](https://doc.log10x.com/run/input/objectStorage/index/#tenxtemplate-filters).

The index output stream produces a list of [Bloom filters](https://en.wikipedia.org/wiki/Bloom_filter){target="\_blank"} for each [indexWriteResolution](https://doc.log10x.com/run/input/objectStorage/index/#indexwriteresolution "index time window resolution") and [indexWriteByteRange](https://doc.log10x.com/run/input/objectStorage/index/#indexwritebyterange "max byte range size to index 'indexReadObject'")
combination of the target blob. [Query](https://doc.log10x.com/run/input/objectStorage/query/ "Query an object storage container for events matching target criteria") inputs utilize these filter objects
to rule out byte ranges where the query criteria are known NOT to match.

For example, if a target blob weighing 10MB contains events whose timestamps
range from the beginning of the hour to 3min later, and [indexWriteResolution](https://doc.log10x.com/run/input/objectStorage/index/#indexwriteresolution "index time window resolution")
is set to `1min` and [indexWriteByteRange](https://doc.log10x.com/run/input/objectStorage/index/#indexwritebyterange "max byte range size to index 'indexReadObject'") is set to `2Mb`, up to 6 ranges
are indexed separately, where the [templateHash](https://doc.log10x.com/api/js/#TenXBaseObject+templateHash "An alphanumeric encoded value of a 64bit hash of the current object's template field.") and [vars](https://doc.log10x.com/api/js/#TenXBaseObject+vars "An array of  variable sequences extracted from the object's text.")
members of each TenXObject are within that range
are added to a list of bloom filters whose accuracy must not fall below this value.
The greater the accuracy, the greater the list of filters is created.

The [query](https://doc.log10x.com/run/input/objectStorage/query/ "Query an object storage container for events matching target criteria") input uses Bloom filters to evaluate whether their corresponding byte ranges contain
target search terms with an accuracy (i.e., chance of false positive) set by these values.
In other words, if this value is 95, there is a 5% chance that a byte range that does NOT
contain target terms is fetched and scanned.


### Advanced

#### :material-menu-right-outline:**`indexWriteTemplateMergeInterval`**

Merge template interval.

|Type|Default|Category|
|---|---|---|
|Number|0|Advanced|

Specifies the interval to wait between template merge operation. Each index operation stores output [TenXTemplate](https://doc.log10x.com/run/template/ "Import joint JSON schemas files to expand events into well-defined TenXObjects.")
objects in the [indexWriteContainer](https://doc.log10x.com/run/input/objectStorage/index/#indexwritecontainer "name of target index container"). Index operations merge templates files into a single file periodically,
with the period interval set by this value.


#### :material-menu-right-outline:**`indexObjectStorageArgs`**

Custom Object storage args.

|Type|Default|Category|
|---|---|---|
|List|\[\]|Advanced|

Custom arguments passed as a map to the constructor of the underlying [object storage](https://doc.log10x.com/run/input/objectStorage/index/#indexobjectstoragename "object storage logical name").
This list is expected to hold pairs of key values (e.g., args: \[key1, value1, key2, value2\]).


### General

#### :material-menu-right-outline:**`indexReadPrintProgress`**

Sets whether this input prints throughput stats to the console.

|Type|Default|Category|
|---|---|---|
|Boolean|false|General|

Sets whether this input prints throughput stats to the console.
This value is commonly used when testing an integration to a remote endpoint.


<br/>:material-github: This module is defined in [index/module.yaml](https://github.com/log-10x/modules/tree/main/pipelines/run/modules/input/objectStorage/index/module.yaml "index/module.yaml"){target="\_blank"}.

