Join our community of builders on Discord!

Worker Alerts (watch)

lightchain-worker watch is a small daemon that runs beside a worker. It posts a message to a Discord webhook when something is wrong with the worker, and another when the problem is over. It only reads. It holds no key, sends no transaction, and never restarts, drains or changes the worker. This guide adds it to a worker set up with Run a Worker on Testnet with the CLI. The watch command is in the CLI from release v0.0.1 on.

What it checks

Every 30 seconds watch runs these checks:
CheckIt alerts whenWhat to do
livenessThe worker's health endpoint, http://127.0.0.1:9101/healthz, does not answerThe service is down or stuck: systemctl status lightchain-worker, then journalctl -u lightchain-worker
ollamaOLLAMA_URL/api/tags does not answerStart Ollama, or fix OLLAMA_URL
rpcThe chain cannot be read through RPC_URL, or the newest block is more than 60 seconds oldCheck the host's network and the RPC endpoint
registeredThe worker address is not registered on-chainlcw status; the env file may point at another keystore
suspendedWorkerRegistry has suspended the worker. The message says when the cooldown endsWait for the cooldown, then lcw reinstate
stakeThe stake is under the on-chain minimumlcw top-up-stake AMOUNT
claimsThree sessions in a row went to other workers, and this worker was already eligible for each one 5 blocks before it was claimedThe worker is not claiming: read its logs
Notes on the checks:
  • While rpc is failing, the four on-chain checks below it are not run. They keep their last result, so an unreadable chain never looks like a recovery.
  • claims runs when SESSION_MANAGER_ADDRESS is set, as it is in the CLI guide. It skips session requests that require a capability.
  • A worker that sends its heartbeat to Redis gets one more check, heartbeat. It alerts when the heartbeat is missing or older than three heartbeat intervals. A worker on the external profile, which is what the CLI guide sets up, sends its heartbeat through the worker gateway and has no such check.
  • lcw reinstate and lcw top-up-stake need a CLI release newer than v0.0.1.

How alerts behave

  • Alert. The first time a check fails, watch posts one message.
  • Repeat. While the check keeps failing, it posts again at most once an hour. The title then reads still failing with how long it has lasted.
  • Recovery. When the check passes again, it posts one message that says how long the problem lasted.
  • Retry. If the webhook cannot be reached, the message is sent again at the next check.
A single failed read of a public RPC is enough for an rpc alert, followed by a recovery 30 seconds later:
CodeTEXT

Step 1: Create a webhook

In Discord, on the channel where alerts should arrive:
  1. Open Edit Channel, then Integrations, then Webhooks. You need the Manage Webhooks permission.
  2. Click New Webhook and give it a name.
  3. Click Copy Webhook URL. It looks like https://discord.com/api/webhooks/<id>/<token>.
Anyone who has the URL can post to the channel, so treat it like a password. watch posts Discord's webhook format: a content line and one embed with a title, a description, a color and a timestamp. Any receiver that accepts that format works.

Step 2: Write the settings file

watch reads the worker's own env file, plus one file of its own that holds the webhook URL. Create it so that only root can read it:
CodeBASH
Check that the host can reach the webhook. Discord answers 204 when it accepts a message:
CodeBASH
These settings are optional. Add them to watch.env to change a default:
SettingDefaultMeaning
WATCH_INTERVAL30sTime between checks
WATCH_COOLDOWN1hShortest gap between repeated alerts for a check that keeps failing
WATCH_MISSED_CLAIMS3Sessions lost in a row before the claims alert; 0 turns the check off
WATCH_WORKER_ADDRESSthe keystore's addressThe worker address to watch. Needs a release newer than v0.0.1

Step 3: Install the service

The unit is a template: the name after @ is the worker's directory under /etc/lightchain/. For the CLI guide's layout that name is worker.
CodeBASH
This line confirms it is running. On the external profile it shows heartbeat=false:
CodeTEXT
The ExecCondition line is a safety check. An older worker-only build ignores its arguments, so started as lightchain-worker watch it would run as a second worker with the same key. With the check, the unit does not start on such a binary. To watch a second worker whose files are in /etc/lightchain/worker-2/, give it its own watch.env and enable lightchain-worker-watch@worker-2.

Troubleshooting

SymptomCause and fix
Failed to load environment filesThe name after @ does not match a directory under /etc/lightchain/, or that directory has no watch.env
watch config invalid … WATCH_WEBHOOK_URL is requiredwatch.env has no WATCH_WEBHOOK_URL line
keystore … has no address fieldThe keystore was not created by init. Add WATCH_WORKER_ADDRESS=0x… to watch.env. On v0.0.1, which lacks that setting, add --worker 0x… to ExecStart in a drop-in instead
systemctl status shows Result: exec-conditionThe binary at /usr/local/bin/lightchain-worker has no watch command. Re-run the installer from the CLI guide
webhook rejected message; retrying next tick with status=404 or 401The webhook was deleted or the URL is wrong. Create a new one and update watch.env
webhook post failed; retrying next tickThe host cannot reach the webhook. Check outbound HTTPS
After changing watch.env, run sudo systemctl restart lightchain-worker-watch@worker.