---
icon: material/kubernetes
---

Deploy the [Retriever](../) to AWS EKS via the [terraform-aws-tenx-retriever](https://registry.terraform.io/modules/log-10x/tenx-retriever/aws){target="\_blank"} module, which provisions EKS workloads, S3, SQS, and IRSA in one `terraform apply`. The underlying [retriever-10x Helm chart](https://github.com/log-10x/helm-charts){target="\_blank"} is cloud-agnostic by design, but non-AWS deployments (GKE, AKS, self-hosted) require S3/SQS endpoint-override wiring that is not yet verified against the current engine build. AWS EKS is the supported and tested path today.

## Runtime and start-up

Retriever pods run the `log10x/quarkus-10x` image, which is a **JVM build**.
No native build of it is published, unlike the Lambda flavour: all 147
published tags are JVM, and `1.1.69-native` returns 404 on both Docker Hub and
GHCR. Size probe budgets for a JVM cold start, not a native one.

Measured on `log10x/quarkus-10x:1.1.68`, three cold container starts on an
8-thread x86_64 host:

| Run | Quarkus `started in` | `/health/started` first returned 200 |
|-----|---------------------|--------------------------------------|
| 1 | 52.2s | 63s |
| 2 | 46.0s | 53s |
| 3 | 35.4s | 44s |

Tens of seconds, not a second or two, and the spread across three runs on one
idle host is already 17 seconds. Faster hosts land lower and contended nodes
land higher. The wall-clock column is the number that matters for Kubernetes,
because it includes container creation and the probe's poll interval.

The `retriever-10x` chart already accounts for this with a **startup probe**
on `/health/started` at `periodSeconds: 10`, `failureThreshold: 30`, a
300-second budget. Kubernetes suspends the liveness and readiness probes until
the startup probe passes, which is why their much shorter
`initialDelaySeconds` (30 and 10) do not restart a pod that is still booting.

!!! warning "Do not delete the startup probe"

    Setting `startupProbe: null` on a cluster, or dropping `failureThreshold`
    below about 12, hands the pod to a liveness probe that fires at 30 seconds
    and restarts a JVM that has not finished starting. The result is a
    crash-loop that looks like a broken image. Keep the startup probe, and
    raise `failureThreshold` rather than lowering it on slow or contended
    nodes.

???+ tenx-bootstrap "Step 1: Prerequisites"

    | Requirement | Description |
    |-------------|-------------|
    | 10x License | Your license key ([get one](https://doc.log10x.com/run/bootstrap/#apikey)) |
    | EKS cluster | With an [OIDC provider](https://docs.aws.amazon.com/eks/latest/userguide/enable-iam-roles-for-service-accounts.html){target="\_blank"} configured for IRSA |
    | AWS IAM | Permissions to create S3 buckets, SQS queues, IAM roles, and S3 event notifications |
    | CLI tools | `kubectl`, `helm`, `terraform` ≥ 1.0, `aws` |
    | GitHub Token | **Optional, for GitOps**: Personal access token for config/symbols repo ([create one](https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/managing-your-personal-access-tokens){target="\_blank"}), see Step 5 |

??? tenx-checklist "Step 2: Deployment Method"

    The [terraform-aws-tenx-retriever](https://registry.terraform.io/modules/log-10x/tenx-retriever/aws){target="\_blank"} module splits configuration into two layers:

    1. **Helm values file**, application settings: clusters, roles per cluster, replicas, resources, autoscaling, Fluent Bit output destination.
    2. **Terraform overrides**, infrastructure values: S3 bucket names, SQS queue URLs, ServiceAccount name. Terraform injects these into the Helm release so you don't wire them by hand.

    See the [Components](#components) section below for the resources Terraform creates.

??? tenx-cloud "Step 3: Configure Infrastructure"

    The module requires your EKS cluster's OIDC provider information for IRSA authentication.

    **Get your EKS OIDC provider info:**

    ```bash
    # Get OIDC provider ARN
    aws eks describe-cluster --name YOUR_CLUSTER_NAME \
      --query "cluster.identity.oidc.issuer" --output text

    # Example output: https://oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE
    # OIDC provider = oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE
    # OIDC provider ARN = arn:aws:iam::ACCOUNT:oidc-provider/oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE
    ```

    **Create Terraform configuration:**

    ``` { .hcl title="main.tf"}
    # Provider configuration - required for the module to access your cluster
    provider "aws" {
      region = "us-east-1"  # Your AWS region
    }

    data "aws_eks_cluster" "cluster" {
      name = "YOUR_CLUSTER_NAME"
    }

    provider "kubernetes" {
      config_path    = "~/.kube/config"
      config_context = "YOUR_KUBECTL_CONTEXT"  # Run: kubectl config current-context
    }

    provider "helm" {
      # Note: 'kubernetes' is an attribute (uses '='), not a block
      kubernetes = {
        config_path    = "~/.kube/config"
        config_context = "YOUR_KUBECTL_CONTEXT"
      }
    }

    # Get current AWS account ID
    data "aws_caller_identity" "current" {}

    # Extract OIDC provider info from EKS cluster
    locals {
      oidc_issuer = replace(data.aws_eks_cluster.cluster.identity[0].oidc[0].issuer, "https://", "")
    }

    module "tenx_retriever" {
      source  = "log-10x/tenx-retriever/aws"

      # Required: API key and OIDC provider for IRSA
      tenx_api_key      = var.tenx_api_key
      oidc_provider_arn = "arn:aws:iam::${data.aws_caller_identity.current.account_id}:oidc-provider/${local.oidc_issuer}"
      oidc_provider     = local.oidc_issuer

      # Kubernetes configuration
      namespace        = "log10x-retriever"
      create_namespace = true

      # Resource naming prefix (used for S3, SQS, IAM resources)
      resource_prefix = "my-app-retriever"

      # S3 bucket configuration (creates new buckets by default)
      # Note: S3 bucket names must be globally unique - uses account ID for uniqueness
      # Set create_s3_buckets = false to use existing buckets
      tenx_retriever_index_source_bucket_name  = "my-app-logs-${data.aws_caller_identity.current.account_id}"
      tenx_retriever_index_results_bucket_name = "my-app-index-${data.aws_caller_identity.current.account_id}"

      # S3 trigger configuration (which files trigger indexing)
      tenx_retriever_index_trigger_prefix = "app/"   # Only index files in app/ prefix
      tenx_retriever_index_trigger_suffix = ".log"   # Only index .log files

      # CloudWatch Logs for query event logging (optional)
      # Creates a log group and grants pods logs:CreateLogStream, logs:PutLogEvents, logs:DescribeLogStreams
      tenx_retriever_query_log_group_name      = "/tenx/prod/retriever/query"
      tenx_retriever_query_log_group_retention = 14   # days (default: 7)

      # Application configuration via Helm values file
      helm_values_file = "${path.module}/values/retriever.yaml"

      # Tagging
      tags = {
        Environment = "production"
        ManagedBy   = "terraform"
      }
    }

    variable "tenx_api_key" {
      description = "10x API key"
      type        = string
      sensitive   = true
    }
    ```

    **Finding your kubectl context:** Run `kubectl config current-context` to get your context name. eksctl creates contexts like `user@cluster.region.eksctl.io`, while `aws eks update-kubeconfig` uses `arn:aws:eks:REGION:ACCOUNT:cluster/NAME`.

    **Using an existing EKS module:** If you're using the [terraform-aws-modules/eks/aws](https://registry.terraform.io/modules/terraform-aws-modules/eks/aws){target="\_blank"} module, replace the OIDC configuration with:

    ```hcl
    oidc_provider_arn = module.eks.oidc_provider_arn
    oidc_provider     = module.eks.oidc_provider
    ```

    **Using existing S3 buckets:** To use existing S3 buckets instead of creating new ones:

    ```hcl
    # Disable bucket creation
    create_s3_buckets = false

    # Reference your existing bucket names
    tenx_retriever_index_source_bucket_name  = "my-existing-logs-bucket"
    tenx_retriever_index_results_bucket_name = "my-existing-logs-bucket"  # Can be same bucket
    ```

    The module will still create the S3→SQS event notification on your existing bucket.

??? tenx-config "Step 4: Configure Deployment Settings"

    Application configuration is provided via a Helm values file. The Terraform module automatically injects infrastructure details (S3 buckets, SQS queues, service account).

    Create a values file for your deployment. Choose a cluster topology:

    === "All-in-one"

        A single cluster handles all roles. Simpler to manage, suitable for development and small workloads.

        ``` { .yaml title="values/retriever.yaml"}
        clusters:
          - name: all-in-one
            roles: ["index", "query", "stream"]
            replicaCount: 2
            maxParallelRequests: 10
            maxQueuedRequests: 1000
            readinessThresholdPercent: 90

            resources:
              requests:
                cpu: 1000m
                memory: 2Gi
              limits:
                memory: 4Gi

            autoscaling:
              enabled: true
              minReplicas: 2
              maxReplicas: 10
              targetCPUUtilizationPercentage: 70

        fluentBit:
          output:
            type: s3  # Options: stdout, s3, cloudwatch, elasticsearch, splunk, datadog
            config:
              s3:
                bucket: my-output-logs-bucket
                region: us-west-2
        ```

        All pods show `2/2 READY` (retriever + fluent-bit sidecar) since every pod includes the `"stream"` role.

    === "Separate clusters per role"

        Dedicated clusters for each role let you scale and size them independently. Each cluster consumes from its own SQS queue:

        - **Index**. Processes S3 event notifications. CPU and memory intensive during file parsing.
        - **Query**. Handles incoming search requests and parallelizes sub-queries across index partitions.
        - **Stream**. Delivers matched events to the Fluent Bit sidecar output.

        ``` { .yaml title="values/retriever.yaml"}
        clusters:
          - name: indexer
            roles: ["index"]
            replicaCount: 2
            maxParallelRequests: 5
            maxQueuedRequests: 500
            readinessThresholdPercent: 90

            resources:
              requests:
                cpu: 2000m
                memory: 4Gi
              limits:
                memory: 6Gi

            autoscaling:
              enabled: true
              minReplicas: 2
              maxReplicas: 10
              targetCPUUtilizationPercentage: 70

          - name: query-handler
            roles: ["query"]
            replicaCount: 3
            maxParallelRequests: 20
            maxQueuedRequests: 1000
            readinessThresholdPercent: 90

            resources:
              requests:
                cpu: 1000m
                memory: 2Gi
              limits:
                memory: 4Gi

            autoscaling:
              enabled: true
              minReplicas: 3
              maxReplicas: 15
              targetCPUUtilizationPercentage: 70

          - name: stream-worker
            roles: ["stream"]
            replicaCount: 5
            maxParallelRequests: 15
            maxQueuedRequests: 1000
            readinessThresholdPercent: 90

            resources:
              requests:
                cpu: 1000m
                memory: 2Gi
              limits:
                memory: 4Gi

            autoscaling:
              enabled: true
              minReplicas: 5
              maxReplicas: 20
              targetCPUUtilizationPercentage: 70

        fluentBit:
          output:
            type: cloudwatch
            config:
              cloudwatch:
                region: us-west-2
                logGroupName: /aws/eks/retriever-logs
                logStreamPrefix: stream-
        ```

        Only `stream-worker` pods show `2/2 READY` (retriever + fluent-bit sidecar). Index and query pods show `1/1`.

    The fluent-bit sidecar is **automatically deployed** for any cluster with `"stream"` in its roles. Without a `fluentBit` section, it defaults to `stdout` output. Add the `fluentBit` section only when you need to override the output destination for production.

    **Note:** The Terraform module creates S3 buckets for log input and index storage, but **not** for Fluent Bit output. If using `type: s3` for Fluent Bit, create the output bucket separately or reference an existing bucket.

    **Values injected automatically by the Terraform module** (do not specify these):

    - `inputBucket` and `indexBucket` (from S3 bucket configuration)
    - `indexQueueUrl`, `queryQueueUrl`, `subQueryQueueUrl`, `streamQueueUrl` (from SQS)
    - `queryLogGroup` (from CloudWatch Logs configuration, sets `TENX_QUERY_LOG_GROUP` env var)
    - `serviceAccount.create = false` and `serviceAccount.name` (Terraform-managed)

    **Scheduled queries** (periodic cron-based querying) are configured separately in Step 11.

??? tenx-githubsync "Step 5: Load Configuration"

    Load the 10x Engine [config folder](https://github.com/log-10x/config){target="\_blank"} into the cluster using one of the methods below.

    If you skip this step, the default configuration bundled with the 10x image is used.

    === ":material-git: Git Repository"

        An init container clones your configuration repository before each pod starts. Works with GitHub, GitLab, Bitbucket, or any HTTPS-accessible Git provider.

        1. Fork the [Config Repository](https://github.com/log-10x/config/fork){target="\_blank"}
        2. Create a branch for your configuration changes
        3. Edit the retriever app configuration in the forked repo

        Add to your Helm values:

        ``` { .yaml title="values/retriever.yaml"}
        config:
          git:
            enabled: true
            url: "https://github.com/YOUR-ACCOUNT/config.git"
            branch: "my-cloud-retriever-config"    # Optional

        symbols:
          git:
            enabled: true
            url: "https://github.com/YOUR-ACCOUNT/symbols.git"
            path: "tenx/my-app/symbols"           # Optional subfolder

        gitToken: "YOUR-GIT-TOKEN"
        ```

        For production, store the token in a Kubernetes Secret rather than in the values file. Consider using [External Secrets Operator](https://external-secrets.io/){target="\_blank"} for automated secret management.

    === ":material-harddisk: Persistent Volume"

        Mount an existing PersistentVolumeClaim that contains your configuration directory. This approach works in air-gapped environments and requires no external network access.

        1. Create a PVC containing your configuration files (cloned from the [Config Repository](https://github.com/log-10x/config){target="\_blank"})
        2. Reference it in your Helm values:

        ``` { .yaml title="values/retriever.yaml"}
        config:
          volume:
            enabled: true
            claimName: "my-config-pvc"

        symbols:
          volume:
            enabled: true
            claimName: "my-symbols-pvc"
        ```

??? tenx-keyfiles "Step 6: Configure Secrets"

    Store sensitive output destination credentials in Kubernetes Secrets.

    **Important**: Only add secrets for outputs you've configured in your [Retriever](https://doc.log10x.com/apps/retriever/run/#configure) app configuration.

    **Create the secret** (example for Elasticsearch):

    ``` { .console .copy }
    kubectl create secret generic cloud-retriever-credentials \
      --from-literal=elastic-username=YOUR_USERNAME \
      --from-literal=elastic-password=YOUR_PASSWORD \
      -n log10x-retriever
    ```

    Add secret references to your Helm values file under each cluster's `extraEnv`:

    ``` { .yaml title="values/retriever.yaml"}
    clusters:
      - name: all-in-one
        roles: ["index", "query", "stream"]
        # ... other config ...

        # Secret references for this cluster
        extraEnv:
          - name: ELASTIC_USERNAME
            valueFrom:
              secretKeyRef:
                name: cloud-retriever-credentials
                key: elastic-username
          - name: ELASTIC_PASSWORD
            valueFrom:
              secretKeyRef:
                name: cloud-retriever-credentials
                key: elastic-password

    # For Datadog output, add to clusters with "stream" role:
    #   - name: DD_API_KEY
    #     valueFrom:
    #       secretKeyRef:
    #         name: cloud-retriever-credentials
    #         key: datadog-api-key

    # For Splunk output, add to clusters with "stream" role:
    #   - name: SPLUNK_HEC_TOKEN
    #     valueFrom:
    #       secretKeyRef:
    #         name: cloud-retriever-credentials
    #         key: splunk-hec-token
    ```

??? tenx-mainconfig "Step 7: Deploy"

    Ensure your AWS credentials are configured (e.g., `AWS_PROFILE=my-profile` or default credentials):

    ``` { .console .copy }
    terraform init
    terraform plan
    terraform apply
    ```

??? tenx-checklist "Step 8: Verify Pods"

    ``` { .console .copy }
    kubectl get pods -n log10x-retriever
    ```

    Pods with the `"stream"` role include a fluent-bit sidecar and show `2/2 READY`. Pods with only `"index"` or `"query"` roles show `1/1 READY`. In the all-in-one configuration (where all roles are combined), all pods show `2/2`.

    **Check retriever console output for startup:**

    The retriever container name matches the cluster name in your values file. Select your topology:

    === "All-in-one"

        ```bash
        kubectl logs -n log10x-retriever -l app=retriever-10x -c retriever-10x-all-in-one --tail=50
        ```

    === "Separate clusters"

        ```bash
        # Check any cluster (indexer, query-handler, or stream-worker)
        kubectl logs -n log10x-retriever -l app=retriever-10x -c retriever-10x-indexer --tail=50
        ```

    Expected startup output:

    ```
    INFO  [com.log10x.ext.quarkus.executor.PipelineExecutor] (main) Initializing cloud accessor...
    INFO  [com.log10x.ext.quarkus.access.aws.AWSAccessor] (main) AWS shared client cache initialized
    INFO  [com.log10x.ext.quarkus.executor.PipelineExecutor] (main) Cloud accessor AWSAccessor initialized
    INFO  [com.log10x.ext.quarkus.executor.PipelineExecutor] (main) Pipeline factory initialized
    INFO  [io.quarkus] (main) run-quarkus 1.1.39 on JVM (powered by Quarkus 3.33.2.1) started in 46.027s. Listening on: http://0.0.0.0:8080
    INFO  [io.quarkus] (main) Installed features: [amazon-sdk-sqs, cdi, rest, rest-jackson, scheduler, smallrye-context-propagation, smallrye-health, vertx]
    ```

    The `started in` figure is a JVM cold start and varies with host CPU and
    node contention, see [Runtime and start-up](#runtime-and-start-up). A pod
    that has not reached `Listening on` yet is still booting, not stuck, until
    the startup probe's 300-second budget runs out.

??? tenx-objectstorageindex "Step 9: Verify Indexing"

    Upload a test file matching your trigger configuration to verify the S3→SQS→indexer pipeline:

    ```bash
    # Create a test log file with current timestamp
    echo "{\"timestamp\":\"$(date -u +%Y-%m-%dT%H:%M:%SZ)\",\"level\":\"ERROR\",\"message\":\"Test error\"}" > test.log

    # Upload to the source bucket (must match trigger prefix/suffix)
    aws s3 cp test.log s3://YOUR-LOGS-BUCKET/app/test.log
    ```

    Check the retriever console output for indexing activity:

    === "All-in-one"

        ```bash
        kubectl logs -n log10x-retriever -l app=retriever-10x -c retriever-10x-all-in-one --tail=50
        ```

    === "Separate clusters"

        ```bash
        kubectl logs -n log10x-retriever -l app=retriever-10x -c retriever-10x-indexer --tail=50
        ```

    Expected indexing output:

    ```
    INFO  [executor-thread-3] IndexTemplates - merged templates. templates.size: 1, rawTemplateFiles.size: 0, ...
    INFO  [executor-thread-3] IndexFilterStats - index complete. Bytes: 21, filters.size: 1, probability historgram: [1=>1], key sizes: [0=>1], element counts: [50=>1], epochs: 1
    INFO  [executor-thread-3] ExecutionPipeline - execution of: /etc/tenx/modules/pipelines/run/pipeline.yaml (myObjectStorageIndex) completed in: 7501ms
    ```

    Verify index results were written to S3:

    ```bash
    aws s3 ls s3://YOUR-INDEX-BUCKET/ --recursive
    ```

    Index files are written to the path configured by `tenx_retriever_index_results_path` (defaults to the bucket root).

    **Monitoring index backlog:**

    If files upload faster than indexing, the SQS queue buffers pending work. Check queue depth and worker status:

    ```bash
    # Check SQS queue depth (number of pending index jobs)
    aws sqs get-queue-attributes \
      --queue-url YOUR-INDEX-QUEUE-URL \
      --attribute-names ApproximateNumberOfMessages

    # Check index worker pod status
    kubectl get pods -n log10x-retriever -l role=index

    # Check autoscaling status
    kubectl get hpa -n log10x-retriever
    ```

??? tenx-objectstoragequery "Step 10: Verify Querying"

    Submit a test query and verify events are streamed to the fluent-bit sidecar. The query JSON format is the same for both methods. The query handler is served at `POST /streamer/query`. The in-engine app path is `@apps/retriever/query`; the public HTTP route keeps the `/streamer/query` name for now, overridable with `LOG10X_RETRIEVER_QUERY_PATH`. Request body fields are `processingTimeMs`, `resultSizeBytes`, and `name`. See [Defining Queries](https://doc.log10x.com/apps/retriever/run/#defining-queries) for the full parameter reference.

    === "HTTP - All-in-one"

        ```bash
        kubectl port-forward -n log10x-retriever svc/tenx-retriever-retriever-10x-all-in-one 8080:80 &

        curl -X POST http://localhost:8080/streamer/query \
          -H "Content-Type: application/json" \
          -d '{"from": "now(\"-5m\")", "to": "now()", "search": "severity_level == \"ERROR\""}'
        ```

        The endpoint returns HTTP 200 to acknowledge receipt.

    === "HTTP - Separate clusters"

        ```bash
        kubectl port-forward -n log10x-retriever svc/tenx-retriever-retriever-10x-query-handler 8080:80 &

        curl -X POST http://localhost:8080/streamer/query \
          -H "Content-Type: application/json" \
          -d '{"from": "now(\"-5m\")", "to": "now()", "search": "severity_level == \"ERROR\""}'
        ```

        The endpoint returns HTTP 200 to acknowledge receipt.

    === "SQS (direct)"

        Send the query JSON directly to the query SQS queue. No port-forward required. Works from any machine with AWS credentials and access to the queue.

        ```bash
        aws sqs send-message \
          --queue-url YOUR-QUERY-QUEUE-URL \
          --message-body '{"from": "now(\"-5m\")", "to": "now()", "search": "severity_level == \"ERROR\""}'
        ```

        Get your query queue URL from the Terraform output:

        ```bash
        terraform output query_queue_url
        ```

    **Check fluent-bit logs for streamed events:**

    Matched events are processed asynchronously via the SQS queues and streamed to the fluent-bit sidecar. Results typically appear within 10-30 seconds in production (up to 60 seconds in single-node setups).

    ```bash
    kubectl logs -n log10x-retriever -l app=retriever-10x -c fluent-bit --tail=50
    ```

    Expected output:

    ```
    [0] com.log10x.ext.cloud.index.query.object.IndexObjectQueryReader0: [[1769076000.000000000, {}], {"timestamp"=>"2026-01-22T10:00:00Z", "level"=>"ERROR", "message"=>"Test error"}]
    ```

    Once running, view your streaming analytics in the [Retriever Dashboard](https://doc.log10x.com/roi-analytics/#cloud-retriever).

??? tenx-config "Step 10b: Track Query Progress"

    When `tenx_retriever_query_log_group_name` is configured (see Step 3), the retriever writes query progress events to a CloudWatch Logs log group. Each query creates log streams under `{queryID}/{workerID}` with scan stats, stream throughput, and error details.

    | Terraform Variable | Description | Default |
    |---|---|---|
    | `tenx_retriever_query_log_group_name` | CloudWatch Logs log group name. If empty, query logging is disabled. | `""` |
    | `tenx_retriever_query_log_group_retention` | Retention in days | `7` |

    The Terraform module creates the log group and grants the IRSA role the necessary permissions (`logs:CreateLogStream`, `logs:PutLogEvents`, `logs:DescribeLogStreams`).

    **Track query progress** using the [Query Console](https://doc.log10x.com/apps/retriever/query/) with `--follow`:

    ```bash
    python3 console.py \
      --search 'severity_level=="ERROR"' --since 5m \
      --bucket YOUR-LOGS-BUCKET --queue-url YOUR-QUERY-QUEUE-URL \
      --log-group "/tenx/prod/retriever/query" \
      --follow
    ```

    The log group can also be set via the `TENX_QUERY_LOG_GROUP` environment variable. The web GUI (`--serve`) displays the configured log group in the Query Log panel.

    **View events directly in CloudWatch:**

    ```bash
    aws logs filter-log-events \
      --log-group-name "/tenx/prod/retriever/query" \
      --start-time $(date -d '5 minutes ago' +%s000) \
      --query-string "ERROR"
    ```

??? tenx-schedule "Step 11: Configure Scheduled Queries"

    The Helm chart can create Kubernetes [CronJobs](https://kubernetes.io/docs/concepts/workloads/controllers/cron-jobs/){target="\_blank"} that periodically send query messages to the SQS query queue. Each job runs an AWS CLI container that submits one or more queries on a cron schedule.

    Add a `scheduledQueries` section to your Helm values file:

    ``` { .yaml title="values/retriever.yaml"}
    scheduledQueries:
      enabled: true
      jobs:
        - name: hourly-errors
          schedule: "0 * * * *"  # Every hour (UTC)
          queries:
            - name: "error-logs"
              from: 'now("-1h")'
              to: 'now()'
              search: 'severity_level=="ERROR"'

        - name: daily-report
          schedule: "0 0 * * *"  # Daily at midnight (UTC)
          queries:
            - name: "daily-errors"
              from: 'now("-24h")'
              to: 'now()'
              search: 'severity_level=="ERROR"'
            - name: "daily-warnings"
              from: 'now("-24h")'
              to: 'now()'
              search: 'severity_level=="WARN"'
    ```

    Each job creates a separate CronJob resource. A single job can contain multiple queries, which are submitted sequentially on each run.

    ### Query Parameters

    | Parameter | Type | Required | Description |
    |-----------|------|----------|-------------|
    | `from` | String | Yes | Start of the time range. Use `now("-1h")`, `now("-24h")`, `now("-7d")`, or a literal epoch. |
    | `to` | String | Yes | End of the time range. Use `now()` for current time. |
    | `search` | String | No | [Search expression](https://doc.log10x.com/run/input/objectStorage/query/#querysearch) for pre-filtering via TenXTemplate filters. |
    | `name` | String | No | Query name. The engine surfaces it as `queryName` in progress logs and summaries. |
    | `processingTimeMs` | Integer (ms) | No | Max processing time per query in milliseconds (e.g., `300000` for 5 minutes). Defaults to 60000 (1 minute). |
    | `resultSizeBytes` | Integer (bytes) | No | Max result volume per query in bytes (e.g., `104857600` for 100 MB). |

    ### Job Options

    | Parameter | Type | Default | Description |
    |-----------|------|---------|-------------|
    | `schedule` | String |, | Cron expression (5 fields: minute hour day month weekday). Runs in UTC. |
    | `suspend` | Boolean | `false` | Temporarily pause a job without deleting it. |
    | `concurrencyPolicy` | String | `Forbid` | Prevents overlapping executions of the same job. |
    | `successfulJobsHistoryLimit` | Integer | `3` | Number of successful job records to keep. |
    | `failedJobsHistoryLimit` | Integer | `1` | Number of failed job records to keep. |
    | `restartPolicy` | String | `OnFailure` | Pod restart policy for failed queries. |

    **Managing scheduled queries:**

    ```bash
    # List all scheduled query CronJobs
    kubectl get cronjobs -n log10x-retriever -l component=scheduled-query

    # View recent job executions
    kubectl get jobs -n log10x-retriever -l component=scheduled-query

    # Manually trigger a scheduled query
    kubectl create job --from=cronjob/tenx-retriever-retriever-10x-hourly-errors manual-run -n log10x-retriever

    # Check logs for a specific job run
    kubectl logs job/manual-run -n log10x-retriever
    ```

    The CronJob pods use the same service account as the retriever, so they automatically inherit SQS permissions. Matched events are delivered to the fluent-bit sidecar output configured in Step 4.

??? tenx-monitoring "Step 12: Monitor Operations"

    ### Health Probes

    Each worker pod exposes [SmallRye Health](https://quarkus.io/guides/smallrye-health){target="_blank"} endpoints with load-based readiness checking. Kubernetes stops routing traffic to pods that reach capacity (configurable via `readinessThresholdPercent`, default 90%):

    | Endpoint | Probe | Description |
    |----------|-------|-------------|
    | `/health/live` | Liveness | Returns `UP` if the process is running |
    | `/health/ready` | Readiness | Returns `READY` below load threshold, `503 NOT READY` at capacity |
    | `/health/started` | Startup | Returns `STARTED` after Quarkus completes initialization |
    | `/metrics/load` |, | JSON metrics: `activeTasks`, `queuedTasks`, `loadPercent`, `totalCapacity` |

    `/health/started` is the one to leave alone. The image is a JVM build, so
    the endpoint stays unanswered for tens of seconds on every cold start (see
    [Runtime and start-up](#runtime-and-start-up)); the chart's defaults of
    `periodSeconds: 10` and `failureThreshold: 30` give it a 300-second budget,
    and Kubernetes holds the liveness and readiness probes off until it passes.
    Raise `failureThreshold` on slow or contended nodes. Lowering it, or
    removing the startup probe, exposes the boot window to the 30-second
    liveness probe and turns a slow start into a crash-loop.

    ### SQS Queue Monitoring

    Monitor queue depth and message age to detect indexing or query backlogs:

    ```bash
    # Check pending messages in index queue
    aws sqs get-queue-attributes \
      --queue-url $LOG10X_INDEX_QUEUE_URL \
      --attribute-names ApproximateNumberOfMessages ApproximateAgeOfOldestMessage
    ```

    **Recommended CloudWatch alarms:**

    - `ApproximateAgeOfOldestMessage` > 300 seconds: indexing or query processing is falling behind
    - `ApproximateNumberOfMessages` growing over time, workers may need scaling

    ### Pod and HPA Status

    ```bash
    # Check pod health and restarts
    kubectl get pods -n log10x-retriever -l app=tenx-retriever

    # Check autoscaler status
    kubectl get hpa -n log10x-retriever
    ```

    ### Recommended: Dead-Letter Queues

    Configure [SQS dead-letter queues](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-dead-letter-queues.html){target="_blank"} with a `RedrivePolicy` on each queue to capture failed messages for inspection instead of silently dropping them after `maxReceiveCount` retries.

??? tenx-delete "Step 13: Teardown"

    To remove all Terraform-managed resources:

    **If the module created S3 buckets** (`create_s3_buckets = true`, the default), you must empty them before destroying. S3 buckets cannot be deleted if they contain objects:

    ```bash
    # Empty buckets first (skip if using create_s3_buckets = false)
    aws s3 rm s3://YOUR-LOGS-BUCKET/ --recursive
    aws s3 rm s3://YOUR-INDEX-BUCKET/ --recursive

    # Destroy all Terraform-managed resources
    terraform destroy
    ```

    When using `create_s3_buckets = false` (see Step 3), Terraform does not manage the S3 buckets, only the S3→SQS event notification is created on them. This means `terraform destroy` will cleanly remove the notification, SQS queues, IAM role, and Kubernetes resources without needing to empty any buckets.

??? tenx-run "Quickstart Full Sample"

    **main.tf** - Complete Terraform configuration:

    ``` { .hcl title="main.tf"}
    provider "aws" {
      region = "us-west-2"
    }

    data "aws_eks_cluster" "cluster" {
      name = "my-production-cluster"
    }

    provider "kubernetes" {
      config_path    = "~/.kube/config"
      config_context = "arn:aws:eks:us-west-2:ACCOUNT:cluster/my-production-cluster"
    }

    provider "helm" {
      # Note: 'kubernetes' is an attribute (uses '='), not a block
      kubernetes = {
        config_path    = "~/.kube/config"
        config_context = "arn:aws:eks:us-west-2:ACCOUNT:cluster/my-production-cluster"
      }
    }

    # Get current AWS account ID
    data "aws_caller_identity" "current" {}

    locals {
      oidc_issuer = replace(data.aws_eks_cluster.cluster.identity[0].oidc[0].issuer, "https://", "")
    }

    module "tenx_retriever" {
      source  = "log-10x/tenx-retriever/aws"

      # Required: API key and OIDC provider
      tenx_api_key      = var.tenx_api_key
      oidc_provider_arn = "arn:aws:iam::${data.aws_caller_identity.current.account_id}:oidc-provider/${local.oidc_issuer}"
      oidc_provider     = local.oidc_issuer

      # Kubernetes
      namespace        = "log10x-retriever"
      create_namespace = true

      # Resource naming prefix
      resource_prefix = "prod-retriever"

      # S3 bucket configuration (names must be globally unique)
      tenx_retriever_index_source_bucket_name  = "prod-logs-${data.aws_caller_identity.current.account_id}"
      tenx_retriever_index_results_bucket_name = "prod-index-${data.aws_caller_identity.current.account_id}"

      # S3 trigger configuration
      tenx_retriever_index_trigger_prefix = "app/"
      tenx_retriever_index_trigger_suffix = ".log"

      # CloudWatch Logs for query event logging (optional)
      tenx_retriever_query_log_group_name      = "/tenx/prod/retriever/query"
      tenx_retriever_query_log_group_retention = 14

      # Application configuration via values file
      helm_values_file   = "${path.module}/values/retriever-production.yaml"

      # Tagging
      tags = {
        Environment = "production"
        Project     = "log10x"
        ManagedBy   = "terraform"
      }
    }

    variable "tenx_api_key" {
      description = "10x API key"
      type        = string
      sensitive   = true
    }
    ```

    **values/retriever-production.yaml** - Application configuration:

    ``` { .yaml title="values/retriever-production.yaml"}
    clusters:
      - name: all-in-one
        roles: ["index", "query", "stream"]
        replicaCount: 2
        maxParallelRequests: 10
        maxQueuedRequests: 1000
        readinessThresholdPercent: 90

        resources:
          requests:
            cpu: 1000m
            memory: 2Gi
          limits:
            memory: 4Gi

        autoscaling:
          enabled: true
          minReplicas: 2
          maxReplicas: 10
          targetCPUUtilizationPercentage: 70

    fluentBit:
      output:
        type: s3
        config:
          s3:
            bucket: my-output-logs-bucket
            region: us-west-2
    ```

    For separate clusters per role (index, query, stream), see the **Separate clusters per role** tab in Step 4.

## :material-cube-outline: Components

What `terraform apply` creates. Tabs describe each component and what it's for.

=== ":material-database-outline: S3 buckets"

    Source and index buckets for the offloaded events. The indexer reads from the source bucket and writes Bloom filter and reverse-index artifacts to the index bucket (which can be the same bucket, the single-bucket layout). Set `create_s3_buckets = false` to reuse buckets you already own.

=== ":material-flash-outline: S3 event notification"

    Triggers indexing. On `ObjectCreated:*` in the source bucket, S3 sends a message to the index SQS queue, which invokes an indexer pod. Scoped by `tenx_retriever_index_trigger_prefix` and `_suffix` so the indexer's own writes under `tenx/` don't re-trigger it.

=== ":material-inbox-multiple-outline: SQS queues"

    Four queues (`index`, `query`, `subquery`, `stream`) buffer work between roles. Queue depth drives HPA-based pod scale-out. A pod crash doesn't drop events in flight. The next poller picks them up.

=== ":material-shield-lock-outline: IAM role (IRSA)"

    The pod identity. Grants least-privilege S3 and SQS access. Assumed via IAM Roles for Service Accounts so no access keys live in the cluster.

=== ":material-account-key-outline: Service Account"

    The Kubernetes ServiceAccount the IRSA role binds to. An annotation on the ServiceAccount links it to the role.

=== ":simple-helm: Helm release"

    Installs the `retriever-10x` chart, which in turn creates pods, HPAs, Services, optional CronJobs for scheduled queries, and the Fluent Bit sidecar spec. Application config (clusters, replicas, resources, Fluent Bit output) lives in your Helm values file.


