Skip to main content
This cookbook explains how to correct data at scale. You write a rule that describes which data points are wrong, and the methods Mangrove should try in order to work out a replacement. The rule then corrects every matching data point in its effective period and keeps correcting new ones as they arrive.

An example scenario

An ambient temperature probe at one site drops out intermittently, so data points arrive as 0. A season of production data is affected, so correcting each value by hand is not practical, and the dropouts will keep happening. You want every affected value replaced with the average of the preceding week, falling back to the last good value when that week has nothing to average.

Two ways to write a correction

Both write to the same correction history on the data point, and both leave the original value readable. Use the single-value endpoint when you know the right answer for a specific data point. Use a substitute rule when the same class of problem recurs.
Substitutions are enabled per account. Ask your Mangrove account team to switch them on. Until they are, the value endpoints return 403 and creating a substitute rule returns 422.

Anatomy of a substitute rule

A substitute rule is an ordinary data rule with two extra fields. It works on data points that already exist on an event, including ones recorded without a value. A data point that ingestion never created at all has no row for a rule to act on, so use the single-value endpoint for those. The condition identifies the data points that need correcting, and the method chain says what to replace them with. Note three constraints on subject_slugs:
  • It takes exactly one slug. A rule corrects one value, even when its condition reads several.
  • The slug has to appear in rule_text.
  • Its data point type has to be numeric.
alert_message is optional on a substitute rule. When you leave it out, Mangrove supplies one, used on the data points the chain could not resolve.

Substitution methods

params: value
Replaces the value with a fixed number. Always resolves, so no method after it can run. Where the condition reads a single value, Mangrove rejects a constant that would itself match it. On a condition reading several values it cannot make that check, so pick the constant with care.
params: slug
Falls back to another data point type (dpt), copying its value from the same event. The target has to be a number type, and defined on the same event type as the value being corrected. Declines when the event carries more than one data point of that type.
no params
Uses the nearest earlier valid value of the same type.
no params
Uses the nearest later valid value of the same type.
params: window
Averages the same type’s values across the window, rounded to five decimal places.
params: window
Takes the lowest value of the same type in the window.
params: window
Takes the highest value of the same type in the window.
params: percentile, window
Takes the chosen percentile of the same type’s values in the window. percentile is a number from 0 to 100, so 90 is the 90th percentile and 50 the median. Where the rank falls between two values, the result is interpolated between them. Declines when the window holds fewer than two values.
The three trailing methods, percentile and both fills draw only on data points that pass the rule and have not themselves been corrected, so a chain never builds a value out of other substitutions. Both fills also decline where several data points tie at the winning timestamp, rather than let the correction depend on the order the data points were recorded in.
They also draw only on data points carrying the same tracking ID as the one being corrected. On a project that captures a tracking ID per delivery or per ticket, that usually leaves no candidates at all, and the rule alerts instead of correcting. If a rule corrects nothing and you expected it to, check whether tracking IDs are set on the events it reads.
Two limits apply. A fill looks back or forward at most 100 data points. A window normally draws on every qualifying data point in it, but when the rule’s condition contains a windowed aggregate or a string comparison, the window is limited to the 1000 data points nearest the one being corrected. trailing_min, trailing_max and the two fills copy an existing value exactly. trailing_average and percentile calculate a new one, rounded to five decimal places.

Windows

The trailing methods and percentile take a window of { kind, value, unit, offset }. Neither kind reaches past the data point being corrected, so a replacement is always derived from data that already existed when the value was captured. A calendar window with no offset is therefore its period to date, not the whole period. backward_fill is the exception: it reads later values. offset steps the window back whole periods, which is how you say “the period before this one” rather than a span trailing the value. It defaults to 0 and takes a non-negative integer. On a calendar window, offset: 1 is the previous bucket, and because that bucket has already closed it is used in full rather than clamped. A monthly average that should draw on last month rather than this one so far is { "kind": "calendar", "unit": "month", "offset": 1 }. A calendar window can span several periods. value counts them and offset places the newest, so { "kind": "calendar", "unit": "year", "offset": 1, "value": 2 } covers the two complete years before the year the data point falls in. With offset: 0 the newest period still ends at the data point itself, and the earlier ones in the span are complete. On a trailing window, an offset shifts the whole span back by its own length. { "kind": "trailing", "value": 7, "unit": "day", "offset": 1 } is the seven days ending a week before the value, not the seven days immediately before it.

How a chain resolves

Methods run in the order you list them and the first one to produce a value wins. A method that cannot produce one is skipped and the next is tried. A value that would match the rule’s own condition is refused, and the chain carries on to the next method. When no method resolves, nothing is written and the data point is flagged with an alert instead. This is the outcome alert_message covers. Mangrove rejects a chain at creation when it repeats a method with the same parameters, or when a method follows a constant.
Two substitute rules matching the same data point correct nothing. Mangrove will not choose between competing rules, so where several fire on one data point it applies none of them and both report the data point as unresolved. The alert names the rules that collided, so start there. Before adding a rule, check that no existing rule already covers the same values, and if a rule’s corrected count comes in lower than you expect, an overlap is the first thing to look for.

Writing the rule

1

Prerequisites

Before starting, make sure you have:
  • A Mangrove API token with Data Rules write permission, and Data Collection read permission to read back what changed. Reverting a value needs Data Collection write as well.
  • Corrections enabled on your account
  • The project ID, and the slug of the data point type you are correcting
2

Create the rule

Post the condition, the subject slug, and the method chain together.
curl
A successful create returns 201 and the saved rule, including the friendly ID you will use to read its results.Mangrove queues an evaluation as soon as the rule is saved, so the corrections land shortly after the call returns rather than during it. A rule with no effective_from or effective_to covers the project’s whole history and everything that arrives afterwards. Set either one to bound it, for example to start at a methodology change.Track a long backfill by re-reading the rule’s stats until evaluated stops climbing.
3

Check what it corrected

The rule’s stats give you the count without paging through every record.
curl
On a substitute rule, corrected counts the data points the chain resolved. measured counts everything else it evaluated, which includes both the data points that passed the condition and the ones the chain could not resolve, so it is not a clean success-versus-failure split.
4

List what it could not correct

On a substitute rule the two result statuses do not mean what they mean on an alert rule, and they do not mean what the app’s filters of the same name mean either.A fired correction is not an alert, so a successful one is recorded as a pass, and so is a proposal still waiting on approval (read approval_state to tell those apart). status=fail is therefore the data points that need attention:
curl
Each of those carries a reason: the chain resolved nothing, the data point sits in a batch in a published report, someone reverted the correction, the write failed, or two substitute rules fired on the same data point.To find the corrections instead, read the results and filter on the per-record corrected flag, which is true where this version wrote a replacement. There is no query parameter for it, so page through and filter client side. For the count alone, use /stats.
The app’s Evaluated Records filter relabels the same two statuses: Substituted maps to fail and Measured to pass, the opposite way round from this endpoint. A count taken from the app will not match a count taken from the API unless you account for that.
A record can be both corrected and fail: the rule corrected it once, a later edit put it back in breach, and a version only ever writes one correction per data point. Widening the window or adding a further method to the chain usually clears most of the unresolved ones. Editing the rule mints a new version and re-evaluates.
5

Inspect one corrected value

Each correction is appended to its data point’s history, which you can read per value.
curl
A rule-written correction comes back with source: rule and the rule named in the rule object, alongside the from_value it replaced and the method that won the chain.

Recipes: stuck sensor and held outlier

The example rule above corrects one value at a time from a fixed condition. The two recipes below use conditions that read a value against its own history, and settings that reach past the single value that fired. Each is a complete request body for Create a rule; swap in slugs that exist in your project. The conditions are explained in Catching a stuck sensor and Spotting an outlier with ZSCORE.

Repair a stuck sensor

The condition is the stuck-sensor check from Rule Expressions: the lowest and highest value over the trailing hour are the same, with at least three values behind it. What the API adds is window_member_scope: repeats, which corrects every repeated value in the window when the condition fires and leaves the first value of the run as measured. The fills and the three trailing methods leave the flat stretch out when they look for a source, so the replacement is never the stuck value itself.
curl
Before creating the rule, run the condition through the preview endpoint. Its response carries suggested_window_member_scope: repeats for a MIN = MAX condition like this one, all for any other windowed condition. Send that value back as window_member_scope. Raise the COUNT to require more identical values before the rule fires. An alert rule with the same condition and window_member_scope: repeats flags the run instead of correcting it.

Replace an outlier and hold it for approval

To catch a value that is inside its permitted range but out of character for the month, test it against the spread of the month’s values. ZSCORE returns how many standard deviations the value sits from the window’s average, so NOT BETWEEN -2 AND 2 matches anything more than two standard deviations out in either direction. It always needs a window, and when the month holds fewer than two values, or they are all identical, it returns nothing and the rule stays quiet. Because an outlier is sometimes real, this rule holds each correction for a person to approve. With requires_approval: true, a match becomes a proposal: the data point keeps its measured value, and nothing is written until someone approves it. The replacement here is the median of the previous complete month, which is what percentile: 50 over a calendar window one period back means.
curl
Deciding the proposals this raises is covered in Deciding proposals over the API.

Deciding proposals over the API

A rule created with requires_approval: true writes nothing on its own. Each match appears in the rule’s results as a pass with approval_state set to pending, and a proposed_correction object carrying the value the rule would write, the method that produced it, and how it arrived at the number. A data point with a pending proposal cannot go into a new batch until someone decides it.
1

Find what is waiting

List the rule’s results filtered to approval_state=pending. The rule’s own pending_count counts what is waiting across all of its active versions.
curl
Each pending entry looks like this, with the derivation trimmed:
2

Approve, or reject with a reason

Approve writes the correction exactly as proposed. Reject needs a reason, which is stored against the result.
curl
Both return the decided result. An approved one carries approval_state: approved and a null proposed_correction; a rejected one carries approval_state: rejected with rejected_at, rejected_by and rejection_reason filled in. A rejection is a different judgment from dismissing an alert, so the dismissed_* fields stay null on it.
To decide up to 100 proposed corrections in one call, send their IDs to bulk approve or bulk reject, with one reason for a bulk reject. Each proposal is decided on its own: the response lists the decided results under data, and each one that could not be decided under errors with its ID and one of the codes below, so a 200 can still carry failures. That includes not_approver: a caller without approval rights gets a 200 with every ID under errors, not a 403. An unknown ID fails the whole call with 404 and decides nothing.
curl
A decision is final. Deciding a result again returns 422 with the code not_pending. The other codes a decision can return: Turning requires_approval off on the rule withdraws every proposal still pending: nobody approved them, so those data points keep their measured values. Turning it on or off does not create a new rule version.

Reading corrected values back

Every data point read carries its correction state, so you do not have to fetch the history to tell whether a value was changed. On the correction itself, origin reads substituted where a value was changed, imputed where one was filled in, and reverted where a correction was undone.
A rule never corrects a data point that has gone into a batch in a published report. Those are skipped and reported as skipped, so a rule run over a period you have already reported changes less than its match count suggests.
A correction does not recalculate anything downstream. A batch that was already generated keeps the value it was generated with. Data rules do re-evaluate against the corrected value, so a correction can clear an alert or raise a new one.

What is only in the app

Two parts of the substitution surface have no API equivalent:
  • Derivation evidence. The stage-by-stage record of which methods were tried and why each was skipped shows in the event drawer. On a written correction the API returns the winning method.
  • Previewing a substitution. The preview endpoint dry-runs the condition and reports which data points would match. It does not compute what the chain would produce.

Undoing a correction

DELETE on the value endpoint restores the original value, the one ingestion delivered, however many corrections stand on it, and appends one revert entry to the history recording the jump from the value in force back to the original.
curl
A revert cannot itself be reverted: once the latest entry is a revert, the value is already the original and a second call returns 422. A correction applied after a revert can be reverted in turn. Reverting a value by hand also stops rules from correcting that value again. To put back every value a rule corrected, call Undo all of a rule’s corrections. It removes the rule’s corrections in the background, so each value reads as Original again rather than Reverted, and returns the number of corrections the rule still holds as corrections_outstanding. A correction inside a reported batch, or one somebody has corrected since, stays in place. The count includes corrections in reported batches, so it is an upper bound on how many will be removed. To see when the undo finishes, list the rule’s runs. The undo shows up as a run with trigger set to on_demand whose state moves from reversing to reversed, and its reversal_summary counts what it removed as corrections_deleted and what it left in place as refused. Nothing else on the run marks it as an undo, so it reads like an evaluation that was reversed. To take back what a single run wrote, reverse the run. When you delete a rule, the values it corrected stay corrected, and each keeps the rule’s name and ID in its history.

Troubleshooting

Validation messages come back prefixed with the field they are about.