Run Standalone Agent Discovery on Kubernetes
Overview
The Arthur ML Engine can run standalone as an agent discovery service. On a schedule, it scans your discovery sources and sends every AI agent it finds to your SIEM or a webhook. It doesn't need the Arthur Platform.
- Configuration: one YAML file lists the sources to scan, how often, and the one destination to send results to. Credentials stay in a Kubernetes Secret, and the file refers to them by name.
- What gets sent: each scan sends one JSON event per agent it found, plus one event saying whether the scan succeeded. The event formats are under What gets sent, below.
- Supported sources and destinations: see the integrations page.
Before you begin
- A Kubernetes cluster with
linux/amd64nodes (the image has no ARM build), andkubectlaccess to it. - The
arthurplatform/ml-engineimage, version2.1.994or later. - Outbound HTTPS from the cluster to your sources and destination.
- A credential for each source and for the destination.
Step 1: Write the config
Discovery finds AI agents by asking systems that already know about them: an MDM for agents installed on laptops, a cloud AI platform for agents deployed there, or a SIEM for agents seen in your logs. Each of these integrations is optional, so configure only the ones you use. Whatever the sources find is sent to one destination, your SIEM or a webhook.
Save your config as discovery.yaml. This example scans Jamf Pro and Vertex AI every 6 hours and sends the results to Splunk HEC:
version: 1
schedule:
interval: 6h
max_concurrent_scans: 2
# Discovery sources: where the engine looks for AI agents. Every supported
# integration is optional; list only the ones you want scanned.
sources:
- name: corp-macs
vendor: jamf_pro
fields:
- {key: base_url, value: https://acme.jamfcloud.com}
- {key: client_id, value: "${JAMF_CLIENT_ID}"} # quote ${...} inside { }
- {key: client_secret, value: "${JAMF_CLIENT_SECRET}"}
configs:
- name: all-macs
query: ""
query_language: none
lookback_window_seconds: 21600
- name: vertex-prod
vendor: gcp_vertex
fields:
- {key: project_id, value: my-gcp-project}
- {key: location, value: us-central1}
- {key: service_account_key, value: "${GCP_SERVICE_ACCOUNT_KEY}"}
configs:
- name: agent-engines
query: ""
query_language: none
lookback_window_seconds: 21600
# Destination: where discovered agents are sent. Exactly one, either Splunk HEC
# (type: splunk_hec) or any HTTPS endpoint (type: webhook).
destination:
type: splunk_hec
url: https://splunk.example.com:8088/services/collector/event
token: ${SPLUNK_HEC_TOKEN}
index: ai_inventoryEach ${NAME} is a credential you'll store in the Secret in Step 2. Every key is described in the Configuration reference below. Load the file into a ConfigMap:
kubectl create namespace arthur-discovery
kubectl -n arthur-discovery create configmap arthur-discovery-config \
--from-file=discovery.yaml --dry-run=client -o yaml | kubectl apply -f -Step 2: Create the Secret
Create one key for each ${NAME} in your config. You need only the credentials for the sources and destination you configured; the keys below match the example. Each key becomes an environment variable in the engine's pod.
kubectl -n arthur-discovery create secret generic arthur-discovery-secrets \
--from-literal=SPLUNK_HEC_TOKEN='<HEC token>' \
--from-literal=JAMF_CLIENT_ID='<Jamf client ID>' \
--from-literal=JAMF_CLIENT_SECRET='<Jamf client secret>' \
--from-file=GCP_SERVICE_ACCOUNT_KEY=./gcp-service-account-key.json \
--dry-run=client -o yaml | kubectl apply -f -Step 3: Deploy the engine
Save this as deployment.yaml, replace <version> with the engine version you're deploying, and apply it.
apiVersion: apps/v1
kind: Deployment
metadata:
name: arthur-discovery
namespace: arthur-discovery
labels:
app: arthur-discovery
spec:
replicas: 1 # always 1: a second replica would scan and send everything twice
strategy:
type: Recreate # never two engines at once, even during an upgrade
selector:
matchLabels:
app: arthur-discovery
template:
metadata:
labels:
app: arthur-discovery
spec:
terminationGracePeriodSeconds: 30 # the engine waits up to 15s for running scans
nodeSelector:
kubernetes.io/arch: amd64
securityContext:
runAsUser: 65532
runAsGroup: 65532
runAsNonRoot: true
seccompProfile:
type: RuntimeDefault
containers:
- name: ml-engine
image: arthurplatform/ml-engine:<version>
env:
- name: ML_ENGINE_DISCOVERY_CONFIG # turns on standalone discovery
value: /etc/arthur/config/discovery.yaml
envFrom:
- secretRef:
name: arthur-discovery-secrets
volumeMounts:
- name: config
mountPath: /etc/arthur/config
readOnly: true
resources:
requests:
cpu: 250m
memory: 1Gi # about 700 MiB in use once running
limits:
cpu: "1"
memory: 2Gi
livenessProbe:
exec: # the health endpoint listens on localhost only
command: ["wget", "-qO", "-", "http://localhost:7492/health"]
initialDelaySeconds: 30
periodSeconds: 30
timeoutSeconds: 10
failureThreshold: 5
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
volumes:
- name: config
configMap:
name: arthur-discovery-configkubectl apply -f deployment.yamlThe pod meets the restricted Pod Security Standard and only makes outbound connections, so it needs no Service.
Step 4: Check that it's working
kubectl -n arthur-discovery logs deployment/arthur-discovery -fEach scan ends with a discovery_scan_outcome log line. Look for "succeeded": true on each one, then check your destination for arthur.discovery.agent events. If the pod is in CrashLoopBackOff, the last log line names the problem; see Troubleshooting.
Configuration reference
Every key discovery.yaml accepts. Unknown keys at the top level, in schedule and in destination stop startup, so a typo there can't be silently ignored.
Top level
| Key | Required | Default | Notes |
|---|---|---|---|
version | yes | Always 1 | |
enabled | no | true | false turns standalone discovery off: the engine ignores the rest of the file and connects to the Arthur Platform instead |
schedule | yes | See below | |
sources | yes | At least one | |
destination | yes | Exactly one | |
emit_scan_outcomes | no | true | Also send one event per scan saying whether it succeeded |
schedule
schedule| Key | Default | Notes |
|---|---|---|
interval | required | How often each config is scanned: 90s, 15m, 6h, 1d, or a number of seconds. At least 1m. The next run starts one interval after the last one started. |
run_on_start | true | Scan everything as soon as the pod starts, instead of waiting one interval |
max_concurrent_scans | 1 | How many configs scan at the same time. When more are due, the most overdue goes first. |
scan_timeout | 6h | A scan running longer than this is asked to stop, and stops blocking other scans. At least 1m. |
sources
sourcesEach source has a name (unique), a vendor, a list of fields, and one or more configs. Each config has a name (unique within its source), a query, a query_language and a lookback_window_seconds.
sources:
- name: <your name for this source>
vendor: jamf_pro | gcp_vertex | splunk_enterprise | elastic_security
fields:
- {key: <field>, value: <value or "${SECRET_NAME}">}
configs:
- name: <your name for this scan>
query: "" # SPL or ES|QL for Splunk and Elastic
query_language: none # spl, esql or none
lookback_window_seconds: 21600 # rounded up to whole hoursAny field value can be ${SECRET_NAME}, or {file: /path/in/the/pod} to read it from a mounted file. The engine treats the vendor's secret fields, marked below, as credentials: it never logs them or sends them to your destination.
The lookback window is how far back each scan reads. Nothing is kept between scans, so set it to at least the interval, or activity between two windows is missed. 0 reads everything the source has.
gcp_vertex: Vertex AI Agent Engine
gcp_vertex: Vertex AI Agent EngineLists every Agent Engine in one project and region. query: "", query_language: none. The lookback window isn't used: every scan lists the whole region.
The service account needs permission to list Agent Engines (roles/aiplatform.viewer covers it).
| Field | Secret | Required | Notes |
|---|---|---|---|
project_id | yes | The project ID, not its number | |
location | no | Region; default us-central1 | |
service_account_key | yes | yes* | Service account JSON key. *Can be left out only if the pod has Google credentials of its own (such as GKE Workload Identity) and the pod sets ARTHUR_ENGINE_GCP_VERTEX_ALLOW_ADC=true. |
jamf_pro: Jamf Pro
jamf_pro: Jamf ProReads the agent inventory that Arthur's endpoint collector writes to each managed Mac. query: "" uses the built-in agent catalog; query_language: none.
The API client's role needs Read Computers, plus Read Smart Computer Groups and Read Static Computer Groups if you set include_groups or exclude_groups.
| Field | Secret | Required | Notes |
|---|---|---|---|
base_url | yes | Your tenant, e.g. https://acme.jamfcloud.com. Must be https://. | |
client_id | yes | yes | Jamf API client |
client_secret | yes | yes | Jamf API client |
include_groups | no | Comma-separated computer group names to scan only | |
exclude_groups | no | Comma-separated computer group names to leave out |
splunk_enterprise: Splunk Enterprise
splunk_enterprise: Splunk EnterpriseRuns your SPL as a search over the lookback window. query_language: spl. The search must return the columns external_id, name and last_seen (use table and rename), and no columns other than those the discovery output format defines.
The token's user needs permission to run searches on the indexes your query reads.
| Field | Secret | Required | Notes |
|---|---|---|---|
base_url | yes | The management port, e.g. https://splunk.example.com:8089 | |
auth_token | yes | yes | A Splunk authentication token |
ca_certificate | no | PEM certificate to trust, in addition to the system CAs | |
tls_verification | no | full (default), ca_only or off; see TLS below |
elastic_security: Elastic Security
elastic_security: Elastic SecurityRuns your ES|QL over the lookback window. query_language: esql (Query DSL isn't supported). Name the output columns with STATS ... BY, EVAL, RENAME and KEEP, to the same three required columns as Splunk.
The API key needs permission to run ES|QL queries on the indices your query reads.
| Field | Secret | Required | Notes |
|---|---|---|---|
elasticsearch_url | yes | Must be https:// | |
api_key | yes | yes | The key's encoded value |
ca_certificate | no | As for Splunk | |
tls_verification | no | As for Splunk |
destination
destinationOne of type: splunk_hec or type: webhook. Both take:
| Key | Default | Notes |
|---|---|---|
url | required | Must be https:// unless allow_insecure_http: true |
tls_verification | full | full, ca_only or off; see TLS below. Quote "off". |
ca_certificate | PEM, inline or {file: path}. Required by ca_only. | |
allow_insecure_http | false | Allow http://. Your token then travels unencrypted. |
batch_size | 100 | Events per request |
timeout_seconds | 30 | Per request |
splunk_hec key | Default | Notes |
|---|---|---|
token | required | HEC token. Must not be empty. |
index | the token's default index | |
source | arthur-ml-engine | |
sourcetype | arthur:discovered_agent |
Point a Splunk HEC url at /services/collector/event.
webhook key | Default | Notes |
|---|---|---|
headers | none | Sent with every request, e.g. Authorization: "Bearer ${WEBHOOK_TOKEN}". Treated as secrets. |
A webhook receives each batch as a POST of a JSON array of events.
Delivery: a request that fails with 429, a 5xx or a connection error is retried up to three times. Any other non-2xx answer fails the batch at once. Redirects are never followed, so point url at the final address.
TLS: full checks the certificate's issuer and hostname. ca_only checks only that your ca_certificate signed it; use this for a certificate that names no hostname the pod uses, such as Splunk's default SplunkServerDefaultCert. off skips the check and logs a warning at startup.
What gets sent
Each scan sends one arthur.discovery.agent event per agent it found, then one arthur.discovery.scan_outcome event (turn these off with emit_scan_outcomes: false). Splunk HEC indexes each as its own event with sourcetype=arthur:discovered_agent. A webhook receives a JSON array of events per request.
{
"event_type": "arthur.discovery.agent",
"schema_version": 1,
"observed_at": "2026-10-09T18:00:00.123456+00:00",
"source": {"id": "6b0f2c1e-…", "name": "vertex-prod", "vendor": "gcp_vertex", "config_name": "agent-engines"},
"agent": {
"external_id": "projects/123456789012/locations/us-central1/reasoningEngines/3765489775063072768",
"name": "personal-assistant",
"last_seen": "2026-10-09T17:58:12+00:00"
}
}agent holds whatever the source reported; depending on the source, that can also include creation_source, platform, runs_on, llm_models, tools, sub_agents and data_sources. Fields the source can't see are left out.
{
"event_type": "arthur.discovery.scan_outcome",
"schema_version": 1,
"observed_at": "2026-10-09T18:00:04.512+00:00",
"source": {"id": "6b0f2c1e-…", "name": "vertex-prod", "vendor": "gcp_vertex", "config_name": "agent-engines"},
"outcome": {"succeeded": true, "records_published": 2, "error": null, "error_code": null, "started_at": "…", "finished_at": "…"}
}- The same agent arrives on every scan. Identify it by
source.idplusagent.external_id. - To alert on a broken source, search for
event_type="arthur.discovery.scan_outcome" outcome.succeeded=false. Theerror_codevalues are listed in Troubleshooting.
Day-to-day operations
The engine reads its config and Secret only at startup, so restart it after changing either.
| Task | How |
|---|---|
| Change the config or a credential | Re-run the Step 1 or Step 2 command, then kubectl -n arthur-discovery rollout restart deployment/arthur-discovery |
| Upgrade | Change the image tag in deployment.yaml, then kubectl apply -f deployment.yaml |
| Pause or resume | kubectl -n arthur-discovery scale deployment/arthur-discovery --replicas=0 (or =1) |
| Remove | kubectl delete namespace arthur-discovery |
- Run one replica. Every engine scans every source, so a second replica sends everything twice. To scan more sources at once, raise
max_concurrent_scans. - Don't pause with
enabled: false. That setting switches the engine to the Arthur Platform. - On shutdown, Jamf, Splunk and Elastic scans stop at a safe point and report
cancelled, Vertex scans finish, and the engine exits within 15 seconds.
To mount credentials as files instead of environment variables, put them in a second Secret, mount it, and refer to each by path:
# deployment.yaml: add to the container's volumeMounts and the pod's volumes
volumeMounts:
- name: secret-files
mountPath: /etc/arthur/secrets
readOnly: true
volumes:
- name: secret-files
secret:
secretName: arthur-discovery-files
# discovery.yaml
- {key: service_account_key, value: {file: /etc/arthur/secrets/gcp-key.json}}Troubleshooting
Start with kubectl -n arthur-discovery logs deployment/arthur-discovery. Startup problems end the log with one line naming the cause. Scan problems show as a discovery_scan_outcome line with "succeeded": false and an error_code.
The pod won't start
| What you see | Cause | Fix |
|---|---|---|
references environment variable(s) that are not set: NAME | The config uses ${NAME} but the Secret has no key NAME | Add the key to the Secret, or fix the spelling in the config. Then restart. |
destination.splunk_hec.token: Value error, must not be empty | The Secret key exists but its value is empty | Set a real value in the Secret |
<path>: Extra inputs are not permitted | A misspelt key, e.g. destination.webhook.heders | Fix the key the message names |
Cannot read discovery config /etc/arthur/config/discovery.yaml: No such file or directory | The ConfigMap's key isn't discovery.yaml, or the mount path differs from ML_ENGINE_DISCOVERY_CONFIG | Create the ConfigMap with --from-file=discovery.yaml and keep both paths as in Step 3 |
this engine has no connector for vendor '...' | A vendor this version doesn't support | Use one of the four vendors above, or upgrade |
Pod stuck in Pending | No amd64 node for the nodeSelector | Add an x86-64 node pool. The image can't run on ARM. |
OOMKilled | The memory limit is too low for a large scan | Raise limits.memory |
A scan fails
error_code | Meaning | What to check |
|---|---|---|
authentication_failed | The source rejected the credential (HTTP 401) | The value in the Secret, and whether it was revoked or rotated |
permission_denied | The credential works but lacks access (HTTP 403) | Jamf API role privileges, the Vertex service account's role, the Splunk or Elastic user's permissions |
not_configured | Something in the source's fields or query is wrong: a missing field, an invalid base_url, a group name that doesn't exist, or query output with the wrong columns | The outcome's error names the field or column |
provider_error | The source returned an error or couldn't be reached | Outbound network access to the source; the outcome's error has the source's own message |
publication_failed | The destination didn't accept the events | See below |
cancelled | The scan was stopped by shutdown or by scan_timeout | If it happens on every scan, raise scan_timeout |
internal_error | An unexpected engine error | Send the log to Arthur support |
Records sent before a failure stay sent; the next scan tries again.
Events don't arrive at the destination
The log has a line like Splunk HEC at <host> refused 2 event(s) with HTTP <status>: <response>.
- 401 or 403: the HEC token or webhook header is wrong or disabled.
- 400
Incorrect index: the HEC token isn't allowed to write to theindexyou set. Allow it on the token, or removeindexto use the token's default. - 3xx (
a redirect, which is not followed): theurlredirects elsewhere. Seturlto the final address. certificate verify failed: the destination's certificate isn't trusted. Setca_certificateto the CA that signed it. For Splunk's default HEC certificate, also settls_verification: ca_only.Could not reach ... after 3 attempt(s): a network problem. Check that the pod has outbound access to the destination's host and port, including any NetworkPolicy.
Updated about 1 hour ago