Skip to main content
Rule expressions are written in a small expression language with familiar comparison and logical operators. Each expression is a condition. When it evaluates to true on a record, the rule acts on that record: an alert rule flags it, a substitute rule corrects it.
This page is the reference for the full set of operators and patterns. For a walkthrough of writing your first rule, see Create a Data Rule.

Variables

Variables reference the data Mangrove evaluates the expression against. They are wrapped in double curly braces and always carry a prefix naming what kind of thing they point at.
Use double quotes around a node name that contains spaces or other special characters, as in the third example. The prefix has to match what the rule runs on. A Data Points rule cannot reference {{node.*}}, and a Batch Calculation rule cannot reference {{data-point.*}}. Mismatches are rejected when you save. What the editor offers follows from that:
  • Data Points rules use data point slugs. These rules apply across the whole project, so autocomplete offers every data point type that is captured on an event, and the rule evaluates against any event whose data points satisfy the condition. One rule reads a single event type: a condition referencing types from two of them is rejected when you save, and autocomplete does not narrow the list to stop you writing one. A calculated data point type, one a model computes rather than an event captures, is not offered, and a condition that names one is rejected when you save. Test a computed value with a Batch Calculation rule and a {{node.<name>}} reference instead.
  • Batch Calculation rules use nodes in the selected model.
Type {{ in the rule editor to open autocomplete. Each suggestion shows the variable name and its description, value type, or unit when available, and inserts the full prefixed form for you.

Comparison operators

The right side takes a literal or a second data point, as in the third example, which is how a mass balance or any other check between two measurements is written without a threshold. Both data points have to be captured on the same event: when one of them is absent from an event, the comparison has nothing to read and the rule does not fire.

Logical operators

Combine conditions with AND, OR, and NOT. Use parentheses to group. NOT binds tightest, then AND, then OR.

Range and set operators

Text operators

String literals use single quotes. MATCHES patterns use Ruby regular expression syntax, and each match attempt is bounded by a one-second timeout, so keep patterns anchored and avoid ones that backtrack heavily. There is no NOT MATCHES operator, and the same goes for the other text operators. Negate them with the NOT prefix instead: NOT {{data-point.lot_code}} matches '[A-Z]{2}-[0-9]+'.

Presence operators

Use these to check whether a value was recorded.
Two cases fall between the two operators, and neither one catches them. A text data point holding an empty string is not missing, because a value was stored, and it is not present either. A boolean holding false behaves the same way. Test for those with = '' and = false instead. A value that was never submitted has no data point for the rule to evaluate, so IS MISSING stays out of its reach.

Aggregates

Aggregate functions roll a set of values into a single number. Use one as the left side of a comparison.
SUM, AVG, MIN, MAX and ZSCORE need a numeric data point type and are rejected when you save otherwise. COUNT works on any type. By default an aggregate covers only the data points on the event being evaluated. To ask a question that spans events, add a window. To compare an aggregate against another measurement, or to hold two measurements to a tolerance band, see Catching drift between two data points with RATIO.

Windowed aggregates

Add a WITHIN clause inside the aggregate’s parentheses to widen the set from one event to many. Use it for rules about how often something happens, or about a total across a period.
Three modes are available. Pick the one that matches the question you are asking.

ROLLING

A sliding window centred on the data point being evaluated, reaching the full duration in both directions. For a value captured at a single point in time, ROLLING 7.days therefore spans 14 days: the seven days before it and the seven after. Where the event carries a start and an end, the window reaches out from both, so the span is the event’s own duration plus the window twice. Use it for proximity and anomaly questions, such as flagging a cluster of data points around one point in time.
Durations take two forms: A ROLLING duration has to be greater than zero. ROLLING 0.days is rejected when you save.

TRAILING

A window that ends at the data point being evaluated and reaches back the duration, so it never reads a value captured later. TRAILING 1.hour on a value captured at 14:00 covers 13:00 to 14:00. Where the event carries a start and an end, the window reaches back from the start and ends at the end. Use it for conditions that only the history up to that point can answer, such as a sensor that has stopped changing.
TRAILING takes the same duration forms as ROLLING, and the duration has to be greater than zero in the same way.

CALENDAR

The calendar bucket the evaluated data point falls in, in your account’s timezone. Use it for cohort and reporting-period questions, such as “no more than one of these per calendar month”.
Units are SECOND, MINUTE, HOUR, DAY, WEEK, MONTH, and YEAR. Weeks start on Monday. The sub-hour units suit duplicate detection at a known cadence, such as one data point per clock minute. Calendar buckets are absolute. February 27 and March 2 sit in different MONTH buckets and never pair, however close together they are. Two data points two seconds apart at 12:00:59 and 12:01:01 fall in different MINUTE buckets for the same reason. Use ROLLING when you care about elapsed time in either direction, TRAILING when only earlier values may count, and CALENDAR when you care about the clock bucket.

What to know before you use one

  • The evaluated data point is in the set when the window names its own type. COUNT(...) = 1 is therefore true for a data point with no neighbours, so write COUNT(...) > 1 to mean “more than just this one”. In a rule that references several types, a value is not a member of a window over a different type.
  • The preview evaluates windowed aggregates. It resolves each window the same way a saved rule does, evaluating once per data point rather than once per event, so what you see in the preview is what the rule will do. Because a windowed condition costs more to resolve, the preview works through it in smaller batches, which means the sampling note in Create a Data Rule appears on a smaller project than an unwindowed rule would need.
  • Windows only work on {{data-point.*}} variables. A WITHIN clause on a {{node.*}} reference is rejected when you save, so Batch Calculation rules cannot use one.
  • A value covering a span of time is in the window only if it overlaps it. A data point that ends exactly as the window opens, or starts exactly as it closes, is left out. For a meter that reports a total every 15 minutes, a CALENDAR HOUR bucket holds the four totals inside that hour and not the one that ended as it began. A value recorded at a single moment on the window’s edge is in it. A substitution method’s window follows the same boundary rule.
  • The rule’s effective dates clamp the window. A CALENDAR MONTH bucket is cut short if the rule only became effective partway through that month.
  • You can use more than one windowed aggregate in a rule. Each resolves its own set independently.
  • A windowed rule can act on the whole window, not only the value that fired. That is a setting on the rule’s Then step rather than part of the expression. See Acting on the whole window.

Catching a stuck sensor

A sensor frozen on one value, a flatline, keeps reporting on schedule, so nothing about any single value looks wrong. The check is that the lowest and highest values over a trailing window are the same, with a count beside it saying how many identical values it takes to call the sensor stuck.
The COUNT is required. MIN = MAX is true of a single value, so without it the rule would fire on the first value after every gap. A MIN = MAX test with no windowed COUNT on the same data point is rejected when you save, and the preview rejects it too. The window sets how long the flatline has to last before the rule fires, and the COUNT stops it firing on a lone value after a gap in the data. Size the count against the meter’s reporting interval: a meter reporting every 15 minutes puts four values in an hour, so COUNT >= 3 is met as soon as the hour is flat. The rule fires on the first value that has a flat hour behind it, and on each later value that does, for as long as the sensor stays stuck. The first value of the run (the run is the stretch of repeated values) is left alone, because the sensor was still working when it reported it. To have the rule flag or correct every repeat in the run rather than only the most recent one, turn on Flag the whole run or Correct the whole run on the Then step. See Acting on the whole window, and repair a stuck sensor with a substitute rule over the API for the same rule written as a request. The COUNT may use a longer window than the MIN and MAX, to count values over a day while testing flatness over an hour. The rule then acts only on the data points inside both windows. Values that alternate, such as 412.6 and 412.7, never trip it, and a flat run of zeros is detected like any other value. Equipment that is legitimately off, such as a flare idle overnight, is also a flat run of zeros, so exclude it with a clause like AND {{data-point.flow_rate}} != 0, or use a window longer than the longest idle stretch you expect.

Spotting an outlier with ZSCORE

ZSCORE returns the z-score of the evaluated value: the number of standard deviations it sits from the average of the window, negative when the value is below the average. It asks how unusual a value is for that data point type. A methane content that is inside its permitted range but two standard deviations away from the month’s typical value is the kind of thing it catches.
ZSCORE always needs a window. Without one it would have only the evaluated value to compare against, so a bare ZSCORE(...) is rejected when you save. Any of the three window modes works. When the window holds fewer than two values, or every value in it is identical, there is no spread to measure. ZSCORE then returns nothing and the comparison is false, so the rule never fires on a lone value or a flat run. A CALENDAR MONTH window is thin in the first days of the month: the evaluated value counts in its own window, so on the second day it is judged against one neighbour. To require a minimum number of values in the window, add a COUNT over the same window, ZSCORE({{data-point.methane_content}} WITHIN CALENDAR MONTH) NOT BETWEEN -2 AND 2 AND COUNT({{data-point.methane_content}} WITHIN CALENDAR MONTH) >= 10, or use WITHIN TRAILING 30.days, which always holds a full month of history. For a substitute rule built on ZSCORE, see replace an outlier and hold it for approval.

Catching drift between two data points with RATIO

RATIO divides one measurement by another, so a rule can test a data point against a second stream or against its own history. Use it to catch two meters that should agree drifting apart, or a week’s average sliding away from the year’s.
Each side is a data point reference or an aggregate over one. A plain number is not accepted on either side. The two sides can carry different windows. A short window over a long one measures drift: it asks whether the recent picture has moved away from the baseline. Both sides must name a numeric data point type, and both must sit on the same event type, since a rule reads one event type. Two meters recorded on separate event types cannot be compared in one rule. A text or true/false type is rejected when you save. RATIO returns nothing when either side has no value, and when the denominator is zero. The comparison is false in both cases, so a rule built on RATIO does not fire on a gap.

Data capture rate

DATA_CAPTURE_RATE asks how much of the data you expected actually arrived, so a rule can flag a period that came in short rather than a value that looks wrong. It reads a fact about the data point type itself rather than aggregating values, so it goes on the left of a comparison.
DATA_CAPTURE_RATE returns a fraction rather than a percentage, so compare against 0.95, not 95. The WITHIN clause is required, because a capture rate with no period behind it means nothing. Only two windows resolve, CALENDAR MONTH and CALENDAR YEAR, and CALENDAR YEAR reads the year to date, which is all that exists before the data year closes. ROLLING windows are not accepted.

When a capture rate rule is rejected

The data point type’s event type must declare a cadence, since the cadence is what supplies the denominator. A rule referencing a data point type whose event type has no cadence will not save. See Data capture rate for declaring one. Once a rule is saved, a period with no capture rate worked out yet reads as no value rather than as zero, and the comparison is false. A rule never fires merely because the figure is missing. A rule is evaluated when data arrives, so this condition reports on a period once the stream delivers into it. To watch a stream that is quiet right now, and to see which slots are missing rather than how many, use the Capture Rate view. Four conditions are rejected when you save, and previewing rejects them too:

Booleans and numbers

  • Booleans are written true and false (case-insensitive).
  • Numbers are written without quotes. Both integers and decimals work: 0, 42, 3.14, 0.005. Negative numbers are allowed.

Common patterns

These examples assume your project has data points with the slugs shown. Substitute slugs that exist in your project. For these patterns as complete rules, grouped by the failure each one catches and carrying the Then-step settings, see Commonly used rules.

Tips for writing good rules

  • Phrase the rule as the failure case. A rule with the expression {{data-point.ph}} not between 6.0 and 8.0 triggers when pH is out of range, which is the alert you want.
  • Validate before saving. Validating previews the matches from the last 90 days. Use it to catch over-eager rules before they generate noise. When nothing matches, a range filter appears with the result so you can look further back.
  • Cover every shape of an empty value. IS MISSING catches a data point with no value stored. It misses a text data point holding an empty string and a boolean holding false, so a rule that has to catch those needs each test spelled out with its own variable: {{data-point.lab_sample_id}} IS MISSING OR {{data-point.lab_sample_id}} = ''.
  • Say why in the rule name. Expressions are terse and a rule carries no separate description field, so the name and the alert message are where you explain what the rule is for.