From raw readings to decisions: field telemetry alerts people act on

How to turn a steady stream of field readings into a handful of alerts worth acting on: thresholds that fit, persistence and hysteresis, alert fatigue, and treating gaps as data.

The live readings on our homepage come from a field node at Yarrabin, NSW, reporting every five minutes over Starlink: air temperature, humidity, dew point, pressure, Starlink latency, solar input, battery state of charge and node uptime. The data comes from a consumer-grade weather station and the solar charge controllers’ Bluetooth telemetry, read by a small ARM computer. It is a typical mix for a remote site, covering weather, link health and power health.

At one reading every five minutes, each of those eight values produces 288 readings a day, or more than two thousand between them. Nobody reads them all, and nobody should have to. The value lies in the few moments when a reading ought to change what someone does. This article is about finding those moments and not burying them.

Start with the decision, not the sensor

For every alert, write down three things before configuring anything: who receives it, what they will do, and how quickly they need to do it. If there is no action, it is not an alert. It is a line in a report.

Working this way sorts alerts into tiers almost automatically:

  • Act now. A frost risk tonight for a frost-sensitive planting, or a battery that will run out before morning. These go to a phone.
  • Act today. A node that has been silent for half an hour, or solar input well below normal for this time of day. These go by message and can wait until someone is free.
  • Know this week. Slow drifts, such as a battery that is recovering a little less each day or latency creeping up. These belong in a weekly digest, not on a phone at 3 am.

Most of what a site produces belongs in the third tier.

Thresholds that fit the site

A fixed threshold, such as “alert if the temperature is below 2°C”, is easy to set up and usually too blunt on its own. A few refinements make a large difference.

Persistence. Require a condition to hold for several consecutive readings before alerting. Consumer sensors produce occasional single-reading spikes, and one odd value is not an event.

Hysteresis. Raise an alert at one level and clear it at another. If a battery alert triggers at 30 per cent, clear it at 40, not 31. Otherwise a value hovering near the line produces a stream of alert, clear, alert.

Rate of change. The trend is often more informative than the level. A steady fall in pressure over a few hours is a classic sign of a change in the weather. Solar input well below what the same hour delivered on recent clear days points to shading, soiling or a fault, whatever the absolute number.

Combined conditions. Frost is the standard example. A low evening dew point on a clear, still night is a strong frost signal, because dry air lets the ground and plants cool further, and frost can form at ground level while a thermometer mounted at standard height still reads a little above zero. An alert that combines falling temperature, a low dew point and settled conditions is far more useful than a temperature threshold alone.

Projection. For power, alert on time remaining rather than a fixed percentage. A battery at 45 per cent that has been losing 10 per cent a night through a week of cloud is more urgent than one at 35 per cent in bright sunshine. A simple straight-line projection of hours to empty is enough to be useful.

Then review. Keep a record of every alert and what was done about it. After a season, any alert that never led to an action gets changed or removed.

Alert fatigue is a design fault

When people receive too many alerts that turn out to need nothing, they stop reading all of them, including the ones that matter. It is well documented in hospitals and in IT operations, and farms are no different. The common causes are thresholds that are too tight, no persistence, duplicates, and alerts about things nobody can fix at that hour.

Duplicates deserve special attention. When a node loses its link, every value behind it goes stale at once, and a naive system sends one alert per value. Group dependent alerts so the root cause produces a single message: node offline, and the stale-data alerts behind it suppressed.

Every alert message should carry enough to act on without opening a dashboard: what, where, the current value and the threshold, how long it has been that way, and a suggested first step. Something like “Ridge node battery 28 per cent (alert below 30), falling about 8 per cent a night, roughly three nights left; consider shedding camera load” gets acted on. “BATT_LOW” gets ignored.

Gaps are data too

A missing reading is not zero, and it is not “the same as last time”. A dashboard that keeps showing the last value it received, frozen for three days, is worse than a blank, because it looks healthy. Show the age of every reading, and grey it out or flag it once it goes stale.

Different kinds of gap point to different causes:

  • One value missing while the others report usually means a sensor fault.
  • Everything missing at once means the node or its link is down.
  • A burst of late readings arriving together means the node kept logging through a link outage and caught up afterwards. The data is fine; the link needs attention.

Node uptime is a quiet but useful signal. If it resets, the node rebooted, whether from a brownout, a watchdog or a fault. A node that reboots around dawn every morning is telling you about its battery.

The system also has to notice silence. The check that raises “no data from this site” must run somewhere other than the site, because a node cannot report its own death, and neither can an alerting service running on it. When the link to a site drops, the hub at the other end is what should notice.

Finally, treat consumer-grade sensors for what they are. They are good at trends and at triggering alerts. They drift, and they are less accurate than research instruments. Compare them with the nearest Bureau of Meteorology station from time to time, and do not report their values with more precision than they have.

Test the whole path

An alert nobody receives is worse than no alert, because everyone assumes it is working. Phone numbers change, emails get filtered, app notifications get switched off. Send a test alert through the entire path on a regular schedule and after every change to the system, and confirm that a person actually received it.

Where Bizix Agritech fits

Our field platform collects readings at the edge, can keep a local record through link outages, and carries the data over SD-WAN to where it can be turned into a small number of alerts with clear owners. The same data feeds the baseline assessment, monitoring and reporting work we deliver for agricultural and research sites. Designed and supported in Australia.