Get Started
Worker Alerts (watch)
lightchain-worker watch is a small daemon that runs beside a worker. It posts a message to a Discord webhook when something is wrong with the worker, and another when the problem is over.
It only reads. It holds no key, sends no transaction, and never restarts, drains or changes the worker.
This guide adds it to a worker set up with Run a Worker on Testnet with the CLI. The watch command is in the CLI from release v0.0.1 on.
What it checks
Every 30 secondswatch runs these checks:
Notes on the checks:
- While
rpcis failing, the four on-chain checks below it are not run. They keep their last result, so an unreadable chain never looks like a recovery. claimsruns whenSESSION_MANAGER_ADDRESSis set, as it is in the CLI guide. It skips session requests that require a capability.- A worker that sends its heartbeat to Redis gets one more check,
heartbeat. It alerts when the heartbeat is missing or older than three heartbeat intervals. A worker on the external profile, which is what the CLI guide sets up, sends its heartbeat through the worker gateway and has no such check. lcw reinstateandlcw top-up-stakeneed a CLI release newer than v0.0.1.
How alerts behave
- Alert. The first time a check fails,
watchposts one message. - Repeat. While the check keeps failing, it posts again at most once an hour. The title then reads
still failingwith how long it has lasted. - Recovery. When the check passes again, it posts one message that says how long the problem lasted.
- Retry. If the webhook cannot be reached, the message is sent again at the next check.
rpc alert, followed by a recovery 30 seconds later:
CodeTEXT
Step 1: Create a webhook
In Discord, on the channel where alerts should arrive:- Open Edit Channel, then Integrations, then Webhooks. You need the Manage Webhooks permission.
- Click New Webhook and give it a name.
- Click Copy Webhook URL. It looks like
https://discord.com/api/webhooks/<id>/<token>.
watch posts Discord's webhook format: a content line and one embed with a title, a description, a color and a timestamp. Any receiver that accepts that format works.
Step 2: Write the settings file
watch reads the worker's own env file, plus one file of its own that holds the webhook URL. Create it so that only root can read it:
CodeBASH
204 when it accepts a message:
CodeBASH
watch.env to change a default:
Step 3: Install the service
The unit is a template: the name after@ is the worker's directory under /etc/lightchain/. For the CLI guide's layout that name is worker.
CodeBASH
heartbeat=false:
CodeTEXT
ExecCondition line is a safety check. An older worker-only build ignores its arguments, so started as lightchain-worker watch it would run as a second worker with the same key. With the check, the unit does not start on such a binary.
To watch a second worker whose files are in /etc/lightchain/worker-2/, give it its own watch.env and enable lightchain-worker-watch@worker-2.
Troubleshooting
After changingwatch.env, run sudo systemctl restart lightchain-worker-watch@worker.
Related guides
- Run a Worker on Testnet with the CLI — the worker setup this page builds on.
- Slashing & Rehabilitation — what the
suspendedandstakealerts mean and how to recover. - Dispatcher-free Mode — how workers claim sessions, which is what the
claimscheck follows.