Adding AI to Log Aggregators

Why?
Even with a very modest Home Lab, my Graylog log aggregator consumes over 1.4M log messages every 24 hours. If I don't check often, over 1,000 of those messages can be errors. Usually, just a few errors duplicated _many_ times. I asked myself the other day "I wonder if Graylog can connect to AI these days to help me summarize, and possibly fix, errors?". Well, yay; I found that the free Graylog Open v7 has a built-in MCP server that just needs to be turned on.
What?
- Graylog
- Any agentic AI that supports MCP servers (OpenClaw, Opencode, Claude, pi-mono, Hermes, whatever)
How?
Install Graylog Open
I used docker. I followed these instructions: https://github.com/Graylog2/docker-compose
Create a Read-Only User
Graylog > System > Users and Teams > Create User
Create an API Token for the Read-Only User
This page is helpful: https://go2docs.graylog.org/current/setting_up_graylog/rest_api_access_tokens.htm
Click More next to the username and choose Edit tokens
Give your token a name and an expiration, then click Create Token
Save the token somewhere safe
Turn on the Graylog MCP Server
Graylog > System > Configurations > MCP
I'm not sure if you are supposed to turn them _all_ on, but I did 🙂
Add the Graylog MCP Server to Opencode (or Other Agentic AI)
First, I converted the token to base64
echo -n "jhhga26ov3g0pvlml22f1redubnfavavin8j7cjn20d20popmp9" | base64
amhoZ2EyNm92M2cwcHZsbWwyMmYxcmVkdWJuZmF2YXZpbjhqN2NqbjIwZDIwcG9wbXA5For opencode, I added this to my ~/opencode.json
cat ~/opencode.json
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"graylog": {
"type": "remote",
"url": "http://192.168.1.11:9000/api/mcp",
"enabled": true,
"headers": {
"Authorization": "Basic amhoZ2EyNm92M2cwcHZsbWwyMmYxcmVkdWJuZmF2YXZpbjhqN2NqbjIwZDIwcG9wbXA5=="
}
}
}
}
Test
opencode run "Please test the graylog mcp server to see if it is working and what its capabilities are"
> build · big-pickle
⚙ graylog_get_system_status Unknown
⚙ graylog_get_current_time Unknown
⚙ graylog_get_formatted_system_status Unknown
⚙ graylog_list_streams Unknown
⚙ graylog_list_inputs Unknown
⚙ graylog_list_index_sets Unknown
I'll test the Graylog MCP server by probing its core capabilities in parallel.
⚙ graylog_list_fields Unknown
⚙ graylog_list_indices Unknown
⚙ graylog_list_resource {"resource_type":"stream"}
⚙ graylog_list_resource {"resource_type":"dashboard"}
⚙ graylog_list_resource {"resource_type":"event_definition"}
Server is reachable. Let me test the search/aggregation and resource capabilities.
⚙ graylog_search_messages {"query":"source:\"server\" AND level:6","fields":["source","level","facility","message"],"limit":3,"range_seconds":3600}
⚙ graylog_aggregate_messages {"query":"*","groupings":[{"field":"facility","limit":5}],"metrics":[{"function":"count"}],"range_seconds":3600}
⚙ graylog_describe_resource {"grn":"grn::::stream:000000000000000000000001"}
Now let me test the search and aggregation capabilities with a real query.
⚙ graylog_search_messages {"query":"*","fields":["source","facility","message"],"limit":2,"range_seconds":300}
Let me confirm message search works with actual data.
The Graylog MCP server is fully working. All 12 tools tested successfully:
**Connection/System**
- `graylog_get_system_status` / `get_formatted_system_status` — Graylog 7.1.7 (Noir), leader node, lifecycle running
- `graylog_get_current_time` — returns UTC time from server
**Data sources**
- `graylog_list_streams` — 4 streams (Default, All events, All system events, tcp-5140)
- `graylog_list_inputs` — 1 input: Syslog TCP on port 5140
- `graylog_list_index_sets` — 3 index sets (default, events, system events; 30-40d rotation, 20 index retention)
- `graylog_list_indices` — 3 active indices, ~1.6M docs in `graylog_0`
**Querying**
- `graylog_list_fields` — schema (source, facility, level, message, timestamps, gl2_* metadata)
- `graylog_search_messages` — returns live Untangle firewall syslog data; supports field selection, paging, time ranges
- `graylog_aggregate_messages` — grouping + metrics work (e.g., 69k user-level messages in last hour, top facilities by count)
**Resources**
- `graylog_list_resource` / `describe_resource` — streams, dashboards (Sources, Dennis), event definitions all resolvable via GRN
Note: search queries scope to streams via IDs/categories and timerange is passed separately (not in query). Everything is live and healthy — data is flowing from the Untangle firewall via syslog TCP 5140.Yay! 🎉
opencode run "Count the most common Graylog error messages over the last 6 hours from stream tcp-5140 where the message includes the word error"
> build · big-pickle
Here are the most common error messages over the last 6 hours (113 total matches) from stream **tcp-5140**:
| Count | Message (normalized) |
|---|---|
| 26 | `sshd-session: error: bind [::]:40037: Address already in use` |
| 26 | `sshd-session: error: channel_setup_fwd_listener_tcpip: cannot listen to port: 40037` |
| 24 | `docker: [NLP]: [ERR] Error processing item..` |
| 7 | `sshd-session: syslogin_perform_logout: logout() returned an error` |
| 7 | `dockerd: level=error msg="[resolver] failed to query external DNS server"` / exec / copy stream |
| 6 | `docker: [v3.1.0.4875] TagLib.CorruptFileException: Error reading page header` |
| 3 | `nginx: [error] open() "...html/<path>" failed (2: No such file or directory)` |
| 3 | `whoopsie.service - Service uploading crash reports to the Ubuntu Error Tracker` |
| 2 | `VideoFileInfoReader: Unable to parse media info from file` |
| 2 | `<error code="429"> Indexer is disabled...` (paired with HttpClient HTTP Error 429) |
The dominant pattern is the recurring **sshd port-forwarding failure on port 40037** (52 of 113 messages), followed by the NLP container's repeated "Error processing item".Oops, forgot to stop that autossh session... 🙂
opencode run "How can I fix the 40037 error"
- sshd_config has GatewayPorts yes (line 2 below), so every remote-forward request binds 0.0.0.0/[::]:40037 on all interfaces.
- cloudflared (PID 3440) proxies imessage.faucher.live → http://192.168.1.11:40037. That port is supposed to be served by an SSH reverse tunnel (-R 40037:...) from another machine.
- Incoming SSH sessions repeatedly request the same remote forward; while a previous tunnel connection is still alive, the new bind fails → your two error messages, logged every time. The session then continues silently without the forward (no ExitOnForwardFailure), so the tunnel drops.
- Right now nothing is bound to 40037 → the tunnel is down, so imessage.faucher.live is likely 502ing.
Fixes (recommended order) 1. On the tunnel client, add to its ~/.ssh/config:
ExitOnForwardFailure yes
ServerAliveInterval 30
ServerAliveCountMax 3
and run it under autossh -N -R 40037:... so it exits on forward failure and cleanly retries — this is the real fix; it stops the silent half-dead sessions. 2. On the server (nvidia-k8s), set GatewayPorts no (it's currently yes). Then sshd binds 40037 on loopback only — if two clients race, the second fails fast instead of lingering. But you must also change the cloudflared ingress origin from 192.168.1.11:40037 to 127.0.0.1:40037, or the proxy won't reach it.
3. Kill any stale listener/holder and confirm a single tunnel instance exists:
ss -tlnp | grep 40037
then restart autossh on the client once so it re-establishes the bind.
Want me to apply the server-side change (GatewayPorts no + cloudflared origin → 127.0.0.1) and reload sshd? I'd leave the client-side autossh change to you on the remote machine.
Thank you
Thank you for reading this post. I hope you found it educational and helpful. I welcome your feedback.