Run Standalone Agent Discovery on Kubernetes

Overview

The Arthur ML Engine can run standalone as an agent discovery service. On a schedule, it scans your discovery sources and sends every AI agent it finds to your SIEM or a webhook. It doesn't need the Arthur Platform.

  • Configuration: one YAML file lists the sources to scan, how often, and the one destination to send results to. Credentials stay in a Kubernetes Secret, and the file refers to them by name.
  • What gets sent: each scan sends one JSON event per agent it found, plus one event saying whether the scan succeeded. The event formats are under What gets sent, below.
  • Supported sources and destinations: see the integrations page.

Before you begin

  • A Kubernetes cluster with linux/amd64 nodes (the image has no ARM build), and kubectl access to it.
  • The arthurplatform/ml-engine image, version 2.1.994 or later.
  • Outbound HTTPS from the cluster to your sources and destination.
  • A credential for each source and for the destination.

Step 1: Write the config

Discovery finds AI agents by asking systems that already know about them: an MDM for agents installed on laptops, a cloud AI platform for agents deployed there, or a SIEM for agents seen in your logs. Each of these integrations is optional, so configure only the ones you use. Whatever the sources find is sent to one destination, your SIEM or a webhook.

Save your config as discovery.yaml. This example scans Jamf Pro and Vertex AI every 6 hours and sends the results to Splunk HEC:

version: 1

schedule:
  interval: 6h
  max_concurrent_scans: 2

# Discovery sources: where the engine looks for AI agents. Every supported
# integration is optional; list only the ones you want scanned.
sources:
  - name: corp-macs
    vendor: jamf_pro
    fields:
      - {key: base_url, value: https://acme.jamfcloud.com}
      - {key: client_id, value: "${JAMF_CLIENT_ID}"}         # quote ${...} inside { }
      - {key: client_secret, value: "${JAMF_CLIENT_SECRET}"}
    configs:
      - name: all-macs
        query: ""
        query_language: none
        lookback_window_seconds: 21600

  - name: vertex-prod
    vendor: gcp_vertex
    fields:
      - {key: project_id, value: my-gcp-project}
      - {key: location, value: us-central1}
      - {key: service_account_key, value: "${GCP_SERVICE_ACCOUNT_KEY}"}
    configs:
      - name: agent-engines
        query: ""
        query_language: none
        lookback_window_seconds: 21600

# Destination: where discovered agents are sent. Exactly one, either Splunk HEC
# (type: splunk_hec) or any HTTPS endpoint (type: webhook).
destination:
  type: splunk_hec
  url: https://splunk.example.com:8088/services/collector/event
  token: ${SPLUNK_HEC_TOKEN}
  index: ai_inventory

Each ${NAME} is a credential you'll store in the Secret in Step 2. Every key is described in the Configuration reference below. Load the file into a ConfigMap:

kubectl create namespace arthur-discovery

kubectl -n arthur-discovery create configmap arthur-discovery-config \
  --from-file=discovery.yaml --dry-run=client -o yaml | kubectl apply -f -

Step 2: Create the Secret

Create one key for each ${NAME} in your config. You need only the credentials for the sources and destination you configured; the keys below match the example. Each key becomes an environment variable in the engine's pod.

kubectl -n arthur-discovery create secret generic arthur-discovery-secrets \
  --from-literal=SPLUNK_HEC_TOKEN='<HEC token>' \
  --from-literal=JAMF_CLIENT_ID='<Jamf client ID>' \
  --from-literal=JAMF_CLIENT_SECRET='<Jamf client secret>' \
  --from-file=GCP_SERVICE_ACCOUNT_KEY=./gcp-service-account-key.json \
  --dry-run=client -o yaml | kubectl apply -f -

Step 3: Deploy the engine

Save this as deployment.yaml, replace <version> with the engine version you're deploying, and apply it.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: arthur-discovery
  namespace: arthur-discovery
  labels:
    app: arthur-discovery
spec:
  replicas: 1                      # always 1: a second replica would scan and send everything twice
  strategy:
    type: Recreate                 # never two engines at once, even during an upgrade
  selector:
    matchLabels:
      app: arthur-discovery
  template:
    metadata:
      labels:
        app: arthur-discovery
    spec:
      terminationGracePeriodSeconds: 30   # the engine waits up to 15s for running scans
      nodeSelector:
        kubernetes.io/arch: amd64
      securityContext:
        runAsUser: 65532
        runAsGroup: 65532
        runAsNonRoot: true
        seccompProfile:
          type: RuntimeDefault
      containers:
        - name: ml-engine
          image: arthurplatform/ml-engine:<version>
          env:
            - name: ML_ENGINE_DISCOVERY_CONFIG   # turns on standalone discovery
              value: /etc/arthur/config/discovery.yaml
          envFrom:
            - secretRef:
                name: arthur-discovery-secrets
          volumeMounts:
            - name: config
              mountPath: /etc/arthur/config
              readOnly: true
          resources:
            requests:
              cpu: 250m
              memory: 1Gi                # about 700 MiB in use once running
            limits:
              cpu: "1"
              memory: 2Gi
          livenessProbe:
            exec:                        # the health endpoint listens on localhost only
              command: ["wget", "-qO", "-", "http://localhost:7492/health"]
            initialDelaySeconds: 30
            periodSeconds: 30
            timeoutSeconds: 10
            failureThreshold: 5
          securityContext:
            allowPrivilegeEscalation: false
            capabilities:
              drop: ["ALL"]
      volumes:
        - name: config
          configMap:
            name: arthur-discovery-config
kubectl apply -f deployment.yaml

The pod meets the restricted Pod Security Standard and only makes outbound connections, so it needs no Service.

Step 4: Check that it's working

kubectl -n arthur-discovery logs deployment/arthur-discovery -f

Each scan ends with a discovery_scan_outcome log line. Look for "succeeded": true on each one, then check your destination for arthur.discovery.agent events. If the pod is in CrashLoopBackOff, the last log line names the problem; see Troubleshooting.

Configuration reference

Every key discovery.yaml accepts. Unknown keys at the top level, in schedule and in destination stop startup, so a typo there can't be silently ignored.

Top level

KeyRequiredDefaultNotes
versionyesAlways 1
enablednotruefalse turns standalone discovery off: the engine ignores the rest of the file and connects to the Arthur Platform instead
scheduleyesSee below
sourcesyesAt least one
destinationyesExactly one
emit_scan_outcomesnotrueAlso send one event per scan saying whether it succeeded

schedule

KeyDefaultNotes
intervalrequiredHow often each config is scanned: 90s, 15m, 6h, 1d, or a number of seconds. At least 1m. The next run starts one interval after the last one started.
run_on_starttrueScan everything as soon as the pod starts, instead of waiting one interval
max_concurrent_scans1How many configs scan at the same time. When more are due, the most overdue goes first.
scan_timeout6hA scan running longer than this is asked to stop, and stops blocking other scans. At least 1m.

sources

Each source has a name (unique), a vendor, a list of fields, and one or more configs. Each config has a name (unique within its source), a query, a query_language and a lookback_window_seconds.

sources:
  - name: <your name for this source>
    vendor: jamf_pro | gcp_vertex | splunk_enterprise | elastic_security
    fields:
      - {key: <field>, value: <value or "${SECRET_NAME}">}
    configs:
      - name: <your name for this scan>
        query: ""                      # SPL or ES|QL for Splunk and Elastic
        query_language: none           # spl, esql or none
        lookback_window_seconds: 21600 # rounded up to whole hours

Any field value can be ${SECRET_NAME}, or {file: /path/in/the/pod} to read it from a mounted file. The engine treats the vendor's secret fields, marked below, as credentials: it never logs them or sends them to your destination.

The lookback window is how far back each scan reads. Nothing is kept between scans, so set it to at least the interval, or activity between two windows is missed. 0 reads everything the source has.

gcp_vertex: Vertex AI Agent Engine

Lists every Agent Engine in one project and region. query: "", query_language: none. The lookback window isn't used: every scan lists the whole region.

The service account needs permission to list Agent Engines (roles/aiplatform.viewer covers it).

FieldSecretRequiredNotes
project_idyesThe project ID, not its number
locationnoRegion; default us-central1
service_account_keyyesyes*Service account JSON key. *Can be left out only if the pod has Google credentials of its own (such as GKE Workload Identity) and the pod sets ARTHUR_ENGINE_GCP_VERTEX_ALLOW_ADC=true.

jamf_pro: Jamf Pro

Reads the agent inventory that Arthur's endpoint collector writes to each managed Mac. query: "" uses the built-in agent catalog; query_language: none.

The API client's role needs Read Computers, plus Read Smart Computer Groups and Read Static Computer Groups if you set include_groups or exclude_groups.

FieldSecretRequiredNotes
base_urlyesYour tenant, e.g. https://acme.jamfcloud.com. Must be https://.
client_idyesyesJamf API client
client_secretyesyesJamf API client
include_groupsnoComma-separated computer group names to scan only
exclude_groupsnoComma-separated computer group names to leave out

splunk_enterprise: Splunk Enterprise

Runs your SPL as a search over the lookback window. query_language: spl. The search must return the columns external_id, name and last_seen (use table and rename), and no columns other than those the discovery output format defines.

The token's user needs permission to run searches on the indexes your query reads.

FieldSecretRequiredNotes
base_urlyesThe management port, e.g. https://splunk.example.com:8089
auth_tokenyesyesA Splunk authentication token
ca_certificatenoPEM certificate to trust, in addition to the system CAs
tls_verificationnofull (default), ca_only or off; see TLS below

elastic_security: Elastic Security

Runs your ES|QL over the lookback window. query_language: esql (Query DSL isn't supported). Name the output columns with STATS ... BY, EVAL, RENAME and KEEP, to the same three required columns as Splunk.

The API key needs permission to run ES|QL queries on the indices your query reads.

FieldSecretRequiredNotes
elasticsearch_urlyesMust be https://
api_keyyesyesThe key's encoded value
ca_certificatenoAs for Splunk
tls_verificationnoAs for Splunk

destination

One of type: splunk_hec or type: webhook. Both take:

KeyDefaultNotes
urlrequiredMust be https:// unless allow_insecure_http: true
tls_verificationfullfull, ca_only or off; see TLS below. Quote "off".
ca_certificatePEM, inline or {file: path}. Required by ca_only.
allow_insecure_httpfalseAllow http://. Your token then travels unencrypted.
batch_size100Events per request
timeout_seconds30Per request
splunk_hec keyDefaultNotes
tokenrequiredHEC token. Must not be empty.
indexthe token's default index
sourcearthur-ml-engine
sourcetypearthur:discovered_agent

Point a Splunk HEC url at /services/collector/event.

webhook keyDefaultNotes
headersnoneSent with every request, e.g. Authorization: "Bearer ${WEBHOOK_TOKEN}". Treated as secrets.

A webhook receives each batch as a POST of a JSON array of events.

Delivery: a request that fails with 429, a 5xx or a connection error is retried up to three times. Any other non-2xx answer fails the batch at once. Redirects are never followed, so point url at the final address.

TLS: full checks the certificate's issuer and hostname. ca_only checks only that your ca_certificate signed it; use this for a certificate that names no hostname the pod uses, such as Splunk's default SplunkServerDefaultCert. off skips the check and logs a warning at startup.

What gets sent

Each scan sends one arthur.discovery.agent event per agent it found, then one arthur.discovery.scan_outcome event (turn these off with emit_scan_outcomes: false). Splunk HEC indexes each as its own event with sourcetype=arthur:discovered_agent. A webhook receives a JSON array of events per request.

{
  "event_type": "arthur.discovery.agent",
  "schema_version": 1,
  "observed_at": "2026-10-09T18:00:00.123456+00:00",
  "source": {"id": "6b0f2c1e-…", "name": "vertex-prod", "vendor": "gcp_vertex", "config_name": "agent-engines"},
  "agent": {
    "external_id": "projects/123456789012/locations/us-central1/reasoningEngines/3765489775063072768",
    "name": "personal-assistant",
    "last_seen": "2026-10-09T17:58:12+00:00"
  }
}

agent holds whatever the source reported; depending on the source, that can also include creation_source, platform, runs_on, llm_models, tools, sub_agents and data_sources. Fields the source can't see are left out.

{
  "event_type": "arthur.discovery.scan_outcome",
  "schema_version": 1,
  "observed_at": "2026-10-09T18:00:04.512+00:00",
  "source": {"id": "6b0f2c1e-…", "name": "vertex-prod", "vendor": "gcp_vertex", "config_name": "agent-engines"},
  "outcome": {"succeeded": true, "records_published": 2, "error": null, "error_code": null, "started_at": "…", "finished_at": "…"}
}
  • The same agent arrives on every scan. Identify it by source.id plus agent.external_id.
  • To alert on a broken source, search for event_type="arthur.discovery.scan_outcome" outcome.succeeded=false. The error_code values are listed in Troubleshooting.

Day-to-day operations

The engine reads its config and Secret only at startup, so restart it after changing either.

TaskHow
Change the config or a credentialRe-run the Step 1 or Step 2 command, then kubectl -n arthur-discovery rollout restart deployment/arthur-discovery
UpgradeChange the image tag in deployment.yaml, then kubectl apply -f deployment.yaml
Pause or resumekubectl -n arthur-discovery scale deployment/arthur-discovery --replicas=0 (or =1)
Removekubectl delete namespace arthur-discovery
  • Run one replica. Every engine scans every source, so a second replica sends everything twice. To scan more sources at once, raise max_concurrent_scans.
  • Don't pause with enabled: false. That setting switches the engine to the Arthur Platform.
  • On shutdown, Jamf, Splunk and Elastic scans stop at a safe point and report cancelled, Vertex scans finish, and the engine exits within 15 seconds.

To mount credentials as files instead of environment variables, put them in a second Secret, mount it, and refer to each by path:

# deployment.yaml: add to the container's volumeMounts and the pod's volumes
          volumeMounts:
            - name: secret-files
              mountPath: /etc/arthur/secrets
              readOnly: true
      volumes:
        - name: secret-files
          secret:
            secretName: arthur-discovery-files

# discovery.yaml
      - {key: service_account_key, value: {file: /etc/arthur/secrets/gcp-key.json}}

Troubleshooting

Start with kubectl -n arthur-discovery logs deployment/arthur-discovery. Startup problems end the log with one line naming the cause. Scan problems show as a discovery_scan_outcome line with "succeeded": false and an error_code.

The pod won't start

What you seeCauseFix
references environment variable(s) that are not set: NAMEThe config uses ${NAME} but the Secret has no key NAMEAdd the key to the Secret, or fix the spelling in the config. Then restart.
destination.splunk_hec.token: Value error, must not be emptyThe Secret key exists but its value is emptySet a real value in the Secret
<path>: Extra inputs are not permittedA misspelt key, e.g. destination.webhook.hedersFix the key the message names
Cannot read discovery config /etc/arthur/config/discovery.yaml: No such file or directoryThe ConfigMap's key isn't discovery.yaml, or the mount path differs from ML_ENGINE_DISCOVERY_CONFIGCreate the ConfigMap with --from-file=discovery.yaml and keep both paths as in Step 3
this engine has no connector for vendor '...'A vendor this version doesn't supportUse one of the four vendors above, or upgrade
Pod stuck in PendingNo amd64 node for the nodeSelectorAdd an x86-64 node pool. The image can't run on ARM.
OOMKilledThe memory limit is too low for a large scanRaise limits.memory

A scan fails

error_codeMeaningWhat to check
authentication_failedThe source rejected the credential (HTTP 401)The value in the Secret, and whether it was revoked or rotated
permission_deniedThe credential works but lacks access (HTTP 403)Jamf API role privileges, the Vertex service account's role, the Splunk or Elastic user's permissions
not_configuredSomething in the source's fields or query is wrong: a missing field, an invalid base_url, a group name that doesn't exist, or query output with the wrong columnsThe outcome's error names the field or column
provider_errorThe source returned an error or couldn't be reachedOutbound network access to the source; the outcome's error has the source's own message
publication_failedThe destination didn't accept the eventsSee below
cancelledThe scan was stopped by shutdown or by scan_timeoutIf it happens on every scan, raise scan_timeout
internal_errorAn unexpected engine errorSend the log to Arthur support

Records sent before a failure stay sent; the next scan tries again.

Events don't arrive at the destination

The log has a line like Splunk HEC at <host> refused 2 event(s) with HTTP <status>: <response>.

  • 401 or 403: the HEC token or webhook header is wrong or disabled.
  • 400 Incorrect index: the HEC token isn't allowed to write to the index you set. Allow it on the token, or remove index to use the token's default.
  • 3xx (a redirect, which is not followed): the url redirects elsewhere. Set url to the final address.
  • certificate verify failed: the destination's certificate isn't trusted. Set ca_certificate to the CA that signed it. For Splunk's default HEC certificate, also set tls_verification: ca_only.
  • Could not reach ... after 3 attempt(s): a network problem. Check that the pod has outbound access to the destination's host and port, including any NetworkPolicy.

Did this page help you?