---
title: "Retriever"
description: "stream offloaded events from your S3 bucket to log analyzer and metric\
  \ outputs on demand"
source: "https://github.com/log-10x/modules/tree/main/apps/retriever/app.yaml"
icon: "material/cloud-arrow-right-outline"

---
The Retriever fetches offloaded logs back on demand. It reads the customer-owned S3 bucket where the Receiver sent the cohort it held back from the analyzer, indexes those events, and returns them when asked, to backfill a metric or replay an old incident.

Re-ingest is customer-driven from `results_location`. There is no vendor S3-to-analyzer shipping layer.

An AI agent runs it through the [log10x MCP](https://doc.log10x.com/apps/mcp/): install with `advise_retriever`, query with `retriever_query` and `retriever_series`. The deploy and query guides cover the same steps by hand.

## :material-cog-transfer-outline: Workflow

The app operates in two phases for efficient data handling: **Index** builds searchable filters when files upload to storage, and **Query** retrieves and streams matching events on-demand to log analyzers and dashboards.

### Index

The [Index](https://doc.log10x.com/run/input/objectStorage/index/ "Index files uploaded to a object storage container (e.g. AWS S3 bucket).") phase executes when log files upload to storage (e.g., S3 bucket) to enable in-place [querying](https://doc.log10x.com/apps/retriever/#query).

<div style="text-align: center;">

```mermaid
graph LR
    A["<div style='font-size: 14px;'>⚡ Trigger</div><div style='font-size: 10px; text-align: center;'>File Upload</div>"] --> B["<div style='font-size: 14px;'>📡 Receive</div><div style='font-size: 10px; text-align: center;'>Read Events</div>"]
    B --> C["<div style='font-size: 14px;'>🔄 Transform</div><div style='font-size: 10px; text-align: center;'>Parse &amp; Structure</div>"]
    C --> D["<div style='font-size: 14px;'>🎁 Enrich</div><div style='font-size: 10px; text-align: center;'>Add Context</div>"]
    D --> E["<div style='font-size: 14px;'>📝 Write</div><div style='font-size: 10px; text-align: center;'>Search Indexes</div>"]
    
    classDef trigger fill:#7c3aed88,stroke:#6d28d9,color:#ffffff,stroke-width:2px,rx:8,ry:8
    classDef receive fill:#9333ea88,stroke:#7c3aed,color:#ffffff,stroke-width:2px,rx:8,ry:8
    classDef transform fill:#2563eb88,stroke:#1d4ed8,color:#ffffff,stroke-width:2px,rx:8,ry:8
    classDef enrich fill:#059669,stroke:#047857,color:#ffffff,stroke-width:2px,rx:8,ry:8
    classDef write fill:#ea580c88,stroke:#c2410c,color:#ffffff,stroke-width:2px,rx:8,ry:8
    
    class A trigger
    class B receive
    class C transform
    class D enrich
    class E write
```

</div>

⚡ **Trigger**: File upload events to object storage send notifications directly to the Index SQS queue for asynchronous processing by index workers

📡 **Receive**: Read log events from the uploaded file in object storage (S3)

🔄 **Transform**: Structured log events into [typed objects](https://doc.log10x.com/run/transform/) with typed fields (severity, timestamp, source, message)

🎁 **Enrich**: Add context via [enrichment rules](https://doc.log10x.com/run/initialize/): geo-IP location, severity classification, k8s metadata, lookup tables

📝 **Write**: Generate lightweight [search indexes](https://doc.log10x.com/run/input/objectStorage/index/#tenxtemplate-filters) that map keywords and fields to specific files, enabling queries to skip most files

### :material-hexagon-multiple-outline: Architecture

=== ":material-check: 10x Index"

    S3 uploads send event notifications to an SQS queue, triggering index workers to generate lightweight [search indexes](https://doc.log10x.com/run/input/objectStorage/index/#tenxtemplate-filters) that map keywords and fields to specific files, enabling queries to skip most files.
    
    <figure markdown="span">
      ![Architecture diagram: S3 uploads trigger SQS notifications to index workers that generate lightweight Bloom filter search indexes mapping keywords and fields to specific files](../../assets/tenx-object-storage-index.png){ align=left }
      <figcaption>:white_check_mark: **Search indexes** enable [querying](#query) events in-place</figcaption>
    </figure>

=== ":material-check: 10x Index + Compact"

    The [Receiver in Compact mode](https://doc.log10x.com/apps/receiver/compact/) losslessly compacts events _before_ they upload, so the stored copy is smaller. Stream queries expand events on the fly for processing and streaming.
    
    <figure markdown="span">
      ![Architecture diagram: Receiver (Compact mode) losslessly compacts events before uploading to S3. Stream queries expand events on the fly](../../assets/tenx-object-storage-index-optimize.png){ align=left }
      <figcaption>:white_check_mark: **Receiver (Compact mode)** [losslessly compacts](https://doc.log10x.com/run/transform/#compact) events before they upload.</figcaption>
    </figure>

### :material-target: Query

Execute queries periodically (e.g., k8s CronJob) or on-demand via the [Console](https://doc.log10x.com/apps/retriever/query/#query-console) to populate log analytics dashboards and alerts (e.g., Splunk, Datadog) with selected events.

<div style="text-align: center;">

```mermaid
graph LR
    A["<div style='font-size: 14px;'>⏰ Trigger</div><div style='font-size: 10px; text-align: center;'>Cron/API Call</div>"] --> B["<div style='font-size: 14px;'>📥 Query</div><div style='font-size: 10px; text-align: center;'>Filter & Fetch Events</div>"]
    B --> C["<div style='font-size: 14px;'>🔄 Transform</div><div style='font-size: 10px; text-align: center;'>Parse &amp; Structure</div>"]
    C --> D["<div style='font-size: 14px;'>🎁 Enrich</div><div style='font-size: 10px; text-align: center;'>Add Context</div>"]
    D --> E["<div style='font-size: 14px;'>🔍 Filter</div><div style='font-size: 10px; text-align: center;'>filters[] Expressions</div>"]
    E --> F["<div style='font-size: 14px;'>📤 Stream</div><div style='font-size: 10px; text-align: center;'>Send to Targets</div>"]
    
    classDef trigger fill:#7c3aed88,stroke:#6d28d9,color:#ffffff,stroke-width:2px,rx:8,ry:8
    classDef query fill:#2563eb88,stroke:#1d4ed8,color:#ffffff,stroke-width:2px,rx:8,ry:8
    classDef transform fill:#059669,stroke:#047857,color:#ffffff,stroke-width:2px,rx:8,ry:8
    classDef enrich fill:#ea580c88,stroke:#c2410c,color:#ffffff,stroke-width:2px,rx:8,ry:8
    classDef filter fill:#2563eb88,stroke:#1d4ed8,color:#ffffff,stroke-width:2px,rx:8,ry:8
    classDef stream fill:#16a34a88,stroke:#15803d,color:#ffffff,stroke-width:2px,rx:8,ry:8
    
    class A trigger
    class B query
    class C transform
    class D enrich
    class E filter
    class F stream
```

</div>

⏰ **Trigger**: Queries initiated via scheduled CronJobs or [Console](https://doc.log10x.com/apps/retriever/query/#query-console) calls

📥 **Query**: Identify and retrieve relevant events from storage by app, timeframe, keywords, or custom criteria

🔄 **Transform**: Parse fetched events into [typed objects](https://doc.log10x.com/run/transform/) with typed fields

🎁 **Enrich**: Add context via [enrichment rules](https://doc.log10x.com/run/initialize/): geo-IP, severity, k8s metadata

🔍 **Filter**: Optional `filters[]` JavaScript expressions narrow the result set in-memory after the scan phase, alongside the Bloom `search` expression and `limit`/`format` options. Common filters drop low-severity noise, remove health-check probes, or keep only grouped stack traces.

📤 **Stream**: Output the selected events to log analyzers (Splunk, Elastic, Datadog) via Fluent Bit, or aggregate into metrics and publish to time-series DBs (Datadog, Prometheus). Horizontal scaling ensures consistent fetch times

### :material-hexagon-multiple-outline: Architecture

=== ":simple-fluentbit: Stream :material-arrow-right-thin: Log Analyzer"

    Enrich, filter with `filters[]`, and stream selected events to log analyzers (e.g., Elastic, Splunk, Datadog) via an embedded [Fluent Bit output](https://doc.log10x.com/run/output/event/fluentbit/) to populate dashboards, queries, and alerts. The `search` expression skips most files at scan time, and `filters[]` expressions narrow the result set in-memory before streaming.
    
    <figure markdown="span">
      ![Architecture diagram: Retriever queries select events from S3, enriches and filters them, then streams via Fluent Bit to log analyzers like Splunk, Elastic, or Datadog](../../assets/tenx-object-storage-query-analyzer.svg){ align=left }
      <figcaption>:white_check_mark: **Stream** selected events from storage to log analyzers</figcaption>
    </figure>

=== ":material-chart-line: Stream :material-arrow-right-thin: Time-series DB"

     Enrich, filter with `filters[]`, aggregate and publish events on-the-fly as metrics to [time-series outputs](https://doc.log10x.com/run/output/metric/) (e.g., Datadog, Prometheus) to populate dashboards, queries and alerts. The `search` expression and `filters[]` expressions select which events contribute to metrics before aggregation. Aggregated metrics are published directly to Datadog's metrics API, Prometheus, or other time-series endpoints.
    
    <figure markdown="span">
      ![Architecture diagram: Retriever queries select events from S3, enriches, filters, and aggregates them into metrics published to time-series databases like Datadog or Prometheus](../../assets/tenx-object-storage-query-TSDB.png){ align=left }
      <figcaption>:white_check_mark: **Stream** aggregated events as metrics to time-series DBs</figcaption>
    </figure>

## :material-shield-check-outline: Infrastructure & Security

Retriever runs entirely within your own AWS or Azure account. No log data leaves your infrastructure.

|Topic|Detail|
|---|---|
|[Runs in your account](https://doc.log10x.com/faq/apps/retriever/)|[Kubernetes](https://doc.log10x.com/apps/retriever/deploy/k8s/) or [Lambda](https://doc.log10x.com/apps/retriever/deploy/lambda/) deployment under your control|
|[No automatic data access](https://doc.log10x.com/faq/apps/retriever/)|You control which events to query and stream|
|[Data stays in your bucket](https://doc.log10x.com/faq/apps/retriever/#what-cloud-storage-services-are-supported)|Index and queries operate only on your S3 or Azure Blob files|
|[Works on AWS and Azure](https://doc.log10x.com/faq/apps/retriever/#what-cloud-storage-services-are-supported)|S3, Azure Blob Storage, and any S3-compatible object storage|
|[Kubernetes Secrets](https://doc.log10x.com/faq/apps/retriever/)|Credentials never stored in config files|

See the [Retriever FAQ](https://doc.log10x.com/faq/apps/retriever/) for complete details on deployment, data access, and security guarantees.

<br/>:material-github: This app is defined in [retriever/app.yaml](https://github.com/log-10x/modules/tree/main/apps/retriever/app.yaml "retriever/app.yaml"){target="\_blank"}.

