Why?

Even with a very modest Home Lab, my Graylog log aggregator consumes over 1.4M log messages every 24 hours. If I don't check often, over 1,000 of those messages can be errors. Usually, just a few errors duplicated _many_ times. I asked myself the other day "I wonder if Graylog can connect to AI these days to help me summarize, and possibly fix, errors?". Well, yay; I found that the free Graylog Open v7 has a built-in MCP server that just needs to be turned on. 

What?

  • Graylog
  • Any agentic AI that supports MCP servers (OpenClaw, Opencode, Claude, pi-mono, Hermes, whatever)

How?

Install Graylog Open

I used docker. I followed these instructions: https://github.com/Graylog2/docker-compose

Create a Read-Only User

Graylog > System > Users and Teams > Create User
 

Create an API Token for the Read-Only User

This page is helpful:  https://go2docs.graylog.org/current/setting_up_graylog/rest_api_access_tokens.htm

Click More next to the username and choose Edit tokens


 

 

Give your token a name and an expiration, then click Create Token

Save the token somewhere safe


Turn on the Graylog MCP Server

Graylog > System > Configurations > MCP

I'm not sure if you are supposed to turn them _all_ on, but I did 🙂
 

Add the Graylog MCP Server to Opencode (or Other Agentic AI)

First, I converted the token to base64
 

echo -n "jhhga26ov3g0pvlml22f1redubnfavavin8j7cjn20d20popmp9" | base64
amhoZ2EyNm92M2cwcHZsbWwyMmYxcmVkdWJuZmF2YXZpbjhqN2NqbjIwZDIwcG9wbXA5

For opencode, I added this to my ~/opencode.json

cat ~/opencode.json 
{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "graylog": {
      "type": "remote",
      "url": "http://192.168.1.11:9000/api/mcp",
      "enabled": true,
      "headers": {
        "Authorization": "Basic amhoZ2EyNm92M2cwcHZsbWwyMmYxcmVkdWJuZmF2YXZpbjhqN2NqbjIwZDIwcG9wbXA5=="
      }
    }
  }
}

Test

opencode run "Please test the graylog mcp server to see if it is working and what its capabilities are"

> build · big-pickle

⚙ graylog_get_system_status Unknown
⚙ graylog_get_current_time Unknown
⚙ graylog_get_formatted_system_status Unknown
⚙ graylog_list_streams Unknown
⚙ graylog_list_inputs Unknown
⚙ graylog_list_index_sets Unknown

I'll test the Graylog MCP server by probing its core capabilities in parallel.

⚙ graylog_list_fields Unknown
⚙ graylog_list_indices Unknown
⚙ graylog_list_resource {"resource_type":"stream"}
⚙ graylog_list_resource {"resource_type":"dashboard"}
⚙ graylog_list_resource {"resource_type":"event_definition"}

Server is reachable. Let me test the search/aggregation and resource capabilities.

⚙ graylog_search_messages {"query":"source:\"server\" AND level:6","fields":["source","level","facility","message"],"limit":3,"range_seconds":3600}
⚙ graylog_aggregate_messages {"query":"*","groupings":[{"field":"facility","limit":5}],"metrics":[{"function":"count"}],"range_seconds":3600}
⚙ graylog_describe_resource {"grn":"grn::::stream:000000000000000000000001"}

Now let me test the search and aggregation capabilities with a real query.

⚙ graylog_search_messages {"query":"*","fields":["source","facility","message"],"limit":2,"range_seconds":300}

Let me confirm message search works with actual data.

The Graylog MCP server is fully working. All 12 tools tested successfully:

**Connection/System**
- `graylog_get_system_status` / `get_formatted_system_status` — Graylog 7.1.7 (Noir), leader node, lifecycle running
- `graylog_get_current_time` — returns UTC time from server

**Data sources**
- `graylog_list_streams` — 4 streams (Default, All events, All system events, tcp-5140)
- `graylog_list_inputs` — 1 input: Syslog TCP on port 5140
- `graylog_list_index_sets` — 3 index sets (default, events, system events; 30-40d rotation, 20 index retention)
- `graylog_list_indices` — 3 active indices, ~1.6M docs in `graylog_0`

**Querying**
- `graylog_list_fields` — schema (source, facility, level, message, timestamps, gl2_* metadata)
- `graylog_search_messages` — returns live Untangle firewall syslog data; supports field selection, paging, time ranges
- `graylog_aggregate_messages` — grouping + metrics work (e.g., 69k user-level messages in last hour, top facilities by count)

**Resources**
- `graylog_list_resource` / `describe_resource` — streams, dashboards (Sources, Dennis), event definitions all resolvable via GRN

Note: search queries scope to streams via IDs/categories and timerange is passed separately (not in query). Everything is live and healthy — data is flowing from the Untangle firewall via syslog TCP 5140.

Yay!  🎉
 

opencode run "Count the most common Graylog error messages over the last 6 hours from stream tcp-5140 where the message includes the word error"

> build · big-pickle

Here are the most common error messages over the last 6 hours (113 total matches) from stream **tcp-5140**:

| Count | Message (normalized) |
|---|---|
| 26 | `sshd-session: error: bind [::]:40037: Address already in use` |
| 26 | `sshd-session: error: channel_setup_fwd_listener_tcpip: cannot listen to port: 40037` |
| 24 | `docker: [NLP]: [ERR] Error processing item..` |
| 7 | `sshd-session: syslogin_perform_logout: logout() returned an error` |
| 7 | `dockerd: level=error msg="[resolver] failed to query external DNS server"` / exec / copy stream |
| 6 | `docker: [v3.1.0.4875] TagLib.CorruptFileException: Error reading page header` |
| 3 | `nginx: [error] open() "...html/<path>" failed (2: No such file or directory)` |
| 3 | `whoopsie.service - Service uploading crash reports to the Ubuntu Error Tracker` |
| 2 | `VideoFileInfoReader: Unable to parse media info from file` |
| 2 | `<error code="429"> Indexer is disabled...` (paired with HttpClient HTTP Error 429) |

The dominant pattern is the recurring **sshd port-forwarding failure on port 40037** (52 of 113 messages), followed by the NLP container's repeated "Error processing item".

Oops, forgot to stop that autossh session... 🙂

opencode run "How can I fix the 40037 error"

- sshd_config has GatewayPorts yes (line 2 below), so every remote-forward request binds 0.0.0.0/[::]:40037 on all interfaces.                       

- cloudflared (PID 3440) proxies imessage.faucher.live → http://192.168.1.11:40037. That port is supposed to be served by an SSH reverse tunnel (-R 40037:...) from another machine.                                                                                                   
- Incoming SSH sessions repeatedly request the same remote forward; while a previous tunnel connection is still alive, the new bind fails → your two error messages, logged every time. The session then continues silently without the forward (no ExitOnForwardFailure), so the tunnel drops.             
- Right now nothing is bound to 40037 → the tunnel is down, so imessage.faucher.live is likely 502ing.                                                                                                               
                                                                                                                                       
Fixes (recommended order)                                                                                                                            1. On the tunnel client, add to its ~/.ssh/config:                                                                                
        ExitOnForwardFailure yes                                                                                                       
        ServerAliveInterval 30                                                                                                         
        ServerAliveCountMax 3                                                                                                          
        and run it under autossh -N -R 40037:... so it exits on forward failure and cleanly retries — this is the real fix; it stops the silent half-dead sessions.                                                                                                                                  2. On the server (nvidia-k8s), set GatewayPorts no (it's currently yes). Then sshd  binds 40037 on loopback only — if two clients race, the second fails fast instead of lingering. But you must also change the cloudflared ingress origin from 192.168.1.11:40037 to 127.0.0.1:40037, or the proxy won't reach it.                                                            
                                                                          3. Kill any stale listener/holder and confirm a single tunnel instance exists:                                                    
        ss -tlnp | grep 40037                                                  
then restart autossh on the client once so it re-establishes the bind.                                                         
Want me to apply the server-side change (GatewayPorts no + cloudflared origin → 127.0.0.1) and reload sshd? I'd leave the client-side autossh change to you on the remote machine. 

Thank you

Thank you for reading this post. I hope you found it educational and helpful. I welcome your feedback.