← Lillie AcademyCourse contentslillieearthintelligence.com

Module 10.5: Case Study, Fire Alerts in Production: From 250,000 False Fires to a Navy Raid

Phase: IV, Production and the job (case study)
Level: Advanced
Estimated time: ~5 hours (about 2 h reading and discussion, 2.5 h lab, 0.5 h write-up)
Prerequisites: Module 10 (active-fire physics, hotspots to events), Module 4 (honest validation), Module 6 (change detection).
Portfolio thread: a small, tested firecheck package that filters geostationary fire detections by confidence, scores recurrence by day, and checks alerts against an independent ground-truth log with coverage reporting.


Why this module

Module 10 taught you how a fire pipeline should work. This module shows what happens when a real one meets real data. It is a true account from Lillie Earth Intelligence's Regional Watch, a service that scans the Niger Delta for illegal crude-oil refining camps from space. In one week (September 2026) the team found three bugs in its geostationary fire feed, each of which produced plausible-looking results, and then tested the fixed system against public reports of Nigerian Navy raids.

Nothing here required a new model. Every fix came from four habits: count what you ingested, read the product specification, match the unit of evidence to the thing you are detecting, and check against ground truth that you did not create. These habits are what hiring managers mean by "production experience", and they are rarely taught.

Learning outcomes

After this module you can:

  1. Explain how a geostationary active-fire product encodes confidence (a probability and a class), and why a threshold must come from the product specification, not from intuition.
  2. Detect silent data loss by reconciling what you downloaded, what you parsed, and what you scored.
  3. Choose a unit of evidence (frames, days, events) that matches the phenomenon, and show how the wrong unit saturates a score.
  4. Validate alerts against an independent ground-truth log: remove region-wide burst days, compare with the base rate at random places, report data coverage, and treat a data gap as "no data", never as a miss.
  5. Design a low-confidence tier that adds sensitivity without contaminating the main alert list.
  6. Write user-facing status that never claims more than the data supports.

Concept lessons (about 2 hours)

Lesson 1: The system and the sensor

The task. Illegal refining camps in the Niger Delta boil stolen crude in open ovens, mostly at night, in mangrove creeks. They are small, intermittent and mobile. Legal gas flares at oil facilities are hot, large and constant. A useful service must find the first and not drown in the second.

The design. The Delta is divided into 226 tiles of 0.25 degrees. For each tile, thermal evidence comes from several sensors:

A hot tile with no mapped facility becomes a refining candidate. A hot tile at a mapped facility is labelled a likely legal flare. The facility map comes from OpenStreetMap and operator registers.

How the product encodes confidence. Each pixel in each 10-minute file carries two fields:

fire_result class fire_probability range Meaning
0 0.00 to 0.19 no fire
1 0.20 to 0.39 low-confidence fire
2 0.40 to 0.79 medium-confidence fire
3 0.80 and above high-confidence fire
4 not set not processed (space, cloud mask, and so on)

The team confirmed this table by opening one real file and cross-tabulating the two fields, not by assuming. Keep that habit.

Lesson 2: Bug one, the pipeline that read 12 files out of 4,400

The collector downloaded a month of files (about 4,400). The scan reported 15 detections. Nobody questioned it, because a quiet rainy season was plausible.

The cause was one line: the parser read only the newest 12 files, a debugging limit left in from early development, when the downloader fetched only the last 6 hours. Twelve 10-minute frames is two hours of data. Every scan for weeks had used about two hours of Meteosat in a 30-day window.

The lesson. Every stage of a pipeline must report its own counts, and someone must compare them:

downloaded 4,414 files  ->  parsed 12 files  ->  15 detections

Written like that, the bug is obvious. Silent sampling is the most common data bug in production remote sensing. The fix was to parse every file, remember each parsed file so reruns are fast, and print "N files, M unreadable, K detections" at the end of every run.

Lesson 3: Bug two, 250,806 fires in one month

With every file parsed, the cache held 250,806 "fires" for September over the Delta, about 56 per 10-minute frame. The scan then marked almost every land tile as a refining candidate.

The sanity check that caught it. When nearly every tile is "hot", the problem is the data, not the Delta. A rate of 56 fires per frame in the rainy season is not physically plausible.

The cause. The parser kept every pixel with fire_probability > 0. In one full-disk frame, about 17,000 pixels of class 0 ("no fire") carry a trace probability of 0.01 to 0.19. Over a month that becomes hundreds of thousands of false detections.

The fix. Parse and store every pixel with probability 0.2 or more, keeping the probability. Apply the alert threshold (0.4, medium and high) when the scan loads the detections, not when the files are parsed. A later threshold change then needs no re-parse of 4,400 files. After the fix: 7,945 stored detections, 1,849 at medium or high confidence in September.

The lesson. "Greater than zero" is not a threshold. Read the product user guide, inspect a real file, and store enough raw information (here, the probability) that you can change your mind later without reprocessing.

Lesson 4: Bug three, counting frames instead of days

The next scan flagged 15 refining candidates. Inspected one by one, 13 had scored the maximum from 2 to 9 frames, mostly on a single day, none at high confidence. Six lay far north of the oil fields, where a one-day fire is most likely farm or bush burning.

The cause. The thermal score saturated at 3 detections. That rule was written when the cache held 12 frames in total. With a 10-minute sensor, one 30-minute grass fire gives 3 frames and a full score.

The fix. Count days with a detection, not frames: 1 day scores 0.33, 2 days 0.67, 3 or more days 1.0. Refining camps burn night after night; a farm fire burns once. Recurrence is the evidence that separates them.

Result: 15 candidates became 6. The strongest had detections on 10 different days and was also seen by MODIS, a second, independent sensor.

The lesson. The unit of evidence must match the behaviour of the target. A rule that was correct for one data rate becomes wrong at another. Whenever you change how much data flows in, re-examine every threshold downstream.

Lesson 5: Validation against ground truth you did not create

A detector that agrees with itself proves nothing. The team needed an independent record of real refining activity, and found one in the news: the Nigerian Navy regularly publishes raids on illegal refining sites, with a date, a Local Government Area (LGA) and sometimes coordinates.

The first raids checked. On 6 September 2026 a Navy unit disrupted two refining sites in Egbema-Ndoni LGA, Rivers State, after an earlier operation in the same area on 28 August. Other September raids followed at Ogbogolo, Aworkiri and Okolomade. The checker looked for Meteosat detections near each raid, from 21 days before to 3 days after.

The first result looked excellent: heat near almost every raid. Then two checks took most of it away.

  1. Burst days. Most "matches" fell on one day, 6 September. That day had 453 detections over the whole Delta, against a normal day of about 20. A day when the entire region lights up (a clear-sky day of farm burning, or a product artefact) says nothing about one site. Such days are now detected and dropped.
  2. Base rate. How often does a random circle of the same size, in the same window, show the same amount of heat? For a 20 km circle, 3 or more heat days happen by chance 21% of the time; for 10 km, 4%.
Raid area Heat days (burst days removed) Chance of that at a random place Reading
Egbema-Ndoni (20 km, LGA only) 3 21% not distinguishable from chance
Okolomade (10 km) 3 4% above chance, weaker source
Aworkiri, Ogbogolo (10 km) 0 not seen

A third check explained why: almost every Meteosat detection over the Delta falls at midday, although files cover all 24 hours. The product sees daytime burning far better than night-time fires, and refining camps often burn at night. For refining, the night-time polar sensor (VIIRS) must lead; Meteosat adds timing.

The lesson. A match is not evidence until you have removed region-wide events and compared against the base rate. The honest result here is "one case above chance out of six checkable raids", which is a starting point, not a claim. Report coverage as well: a window with no data is "no data", never a miss.

Keep the raid log as a test set: if you also train on it, your test score means nothing.

Lesson 6: A weak tier, and status that tells the truth

The weak tier. Lowering the main threshold to 0.2 would bring back the false-alarm flood. Instead, add a separate tier: a tile with no other signal, no mapped facility, and low-confidence heat on 3 or more days becomes a "weak candidate". It has its own colour on the map, a low score (0.2), and it is not counted with the main candidates. Analysts see it; nobody mistakes it for a strong lead.

Status that tells the truth. Two smaller bugs in the same week were about wording, and they matter as much as the data bugs:

"We looked and found nothing" and "we have not looked" must never look the same to a user.


Guided lab (about 2.5 hours): the firecheck package

You will build three tested helpers and run them on synthetic detections that reproduce the three bugs. No download is needed; the generator creates the data.

Step 1: Generate synthetic detections

# firecheck/synth.py
import random
from datetime import date, timedelta


def synth_month(seed=7, start=date(2026, 9, 1), days=30):
    """Detections with the same structure as the MTG fire product.

    - background: trace probabilities (class 0, p < 0.2) everywhere
    - a camp: p 0.25 to 0.6 at one spot on 8 random nights
    - a farm fire: p 0.5, six frames in one afternoon, elsewhere
    """
    rng = random.Random(seed)
    out = []
    for d in range(days):
        day = start + timedelta(days=d)
        for _ in range(300):                       # trace noise
            out.append({"lat": rng.uniform(4.5, 6.5), "lon": rng.uniform(5.0, 7.5),
                        "p": round(rng.uniform(0.01, 0.19), 2), "time": f"{day}T12:00Z"})
    for d in rng.sample(range(days), 8):           # recurring camp
        day = start + timedelta(days=d)
        out.append({"lat": 4.80, "lon": 6.10, "p": round(rng.uniform(0.25, 0.6), 2),
                    "time": f"{day}T22:00Z"})
    for m in range(0, 60, 10):                     # one farm fire, one day
        out.append({"lat": 6.60, "lon": 5.10, "p": 0.5, "time": f"{start}T13:{m:02d}Z"})
    return out

Step 2: Write the helpers and their tests

# firecheck/core.py
import math
from datetime import date, timedelta


def filter_confidence(dets, p_min=0.4):
    """Keep detections at or above p_min. Detections without p are dropped."""
    return [d for d in dets if d.get("p") is not None and d["p"] >= p_min]


def day_score(dets, box):
    """Thermal score from DISTINCT DAYS in a box: 1 day 0.33, 2 days 0.67, 3+ days 1.0."""
    lat0, lat1, lon0, lon1 = box
    days = {d["time"][:10] for d in dets
            if lat0 <= d["lat"] < lat1 and lon0 <= d["lon"] < lon1}
    return min(1.0, len(days) / 3.0)


def km(lat1, lon1, lat2, lon2):
    return math.hypot((lon2 - lon1) * 111.32 * math.cos(math.radians((lat1 + lat2) / 2)),
                      (lat2 - lat1) * 110.57)


def check_event(event, dets, covered_days, before=21, after=3, radius_km=20):
    """Detections near a ground-truth event, by tier, with coverage.

    covered_days: the set of dates for which ANY data file exists.
    """
    d0 = date.fromisoformat(event["date"])
    window = [(d0 + timedelta(days=k)).isoformat() for k in range(-before, after + 1)]
    cov = sum(day in covered_days for day in window) / len(window)
    if cov == 0:
        return {"status": "no data"}
    near = [d for d in dets if d["time"][:10] in window
            and km(event["lat"], event["lon"], d["lat"], d["lon"]) <= radius_km]
    return {"status": "checked", "coverage": round(cov, 2),
            "medium_days": sorted({d["time"][:10] for d in near if d["p"] >= 0.4}),
            "low_days": sorted({d["time"][:10] for d in near if 0.2 <= d["p"] < 0.4})}
# tests/test_firecheck.py
from firecheck.core import check_event, day_score, filter_confidence
from firecheck.synth import synth_month


def test_threshold_removes_trace_noise():
    dets = synth_month()
    assert len([d for d in dets if d["p"] > 0]) > 8000      # the "p > 0" bug
    assert len(filter_confidence(dets)) < 20                # the fix


def test_farm_fire_does_not_saturate_but_camp_does():
    dets = filter_confidence(synth_month(), 0.2)
    farm = (6.5, 6.75, 5.0, 5.25)
    camp = (4.75, 5.0, 6.0, 6.25)
    assert day_score(dets, farm) < 0.5
    assert day_score(dets, camp) == 1.0


def test_gap_is_no_data_not_a_miss():
    event = {"date": "2026-08-28", "lat": 4.8, "lon": 6.1}
    assert check_event(event, [], covered_days=set())["status"] == "no data"

Step 3: Reproduce the bugs, then fix them

  1. Score every tile with frames instead of days, and with p > 0 instead of p >= 0.4. Count how many tiles become "candidates". Write the number down.
  2. Apply the fixes one at a time and record how the count changes after each. This is your before-and-after table.
  3. Create a small ground-truth log of three fake events: one at the camp, one at the farm fire, and one in a week you delete from the data. Run check_event and confirm the third reports "no data".

Step 4: Write-up and commit

Write a one-page incident note in the style of a production postmortem: what users saw, the cause, how it was detected, the fix, and the check that now prevents it. Commit the package, tests and note.


Checkpoint

Self-check (answers below).

  1. A collector reports 4,400 files downloaded; the scan reports 15 detections for a month. What do you check first?
  2. Why is fire_probability > 0 a bad threshold for this product, and what should decide the threshold?
  3. One 30-minute fire gives three detections from a 10-minute geostationary sensor. Why does that break a rule that saturates at three detections, and what unit fixes it?
  4. Your checker says a raid was "not detected". What must you know before you believe it?
  5. Why should a raid log be kept mainly as a test set, not training data?
  6. Four raids in different places all show heat on the same day. What do you check before calling that a detection?

Interview-style questions (practice out loud).

Answers. (1) Reconcile the counts at each stage: downloaded, parsed, detections. A gap between downloaded and parsed points to a parsing limit or failures.
(2) Class 0 ("no fire") pixels carry trace probabilities, so "greater than zero" includes non-fires by design. The product specification and an inspection of real files decide the threshold, with the raw probability stored so it can change later.
(3) The rule was written for a much lower data rate. Counting distinct days measures recurrence, which separates repeated refining from one-off burning.
(4) The data coverage over the event window. A gap in the data is "no data", not a miss.
(5) A model trained and tested on the same events will score itself well and prove nothing. Keep an independent test set, ideally the most recent events.
(6) Whether that day was a region-wide burst (count the detections over the whole region against a normal day), and the base rate: how often a random place of the same size shows the same heat in the same window.


Deliverable

A firecheck package with the three helpers, the synthetic generator, passing tests, a before-and-after table showing the effect of each fix, and a one-page incident note.

What a hiring manager sees

Anyone can train a model. Few candidates can show they found and fixed silent data loss, a wrong threshold and a saturating score in a live system, then validated it against independent ground truth with honest coverage reporting. This module gives you a real story to tell, and a small package that proves you can tell it with tests.

Currency note

Verified October 2026. The MTG FCI active-fire product is distributed by EUMETSAT through its Data Store (collection EO:EUM:DAT:0682), one full-disk file every 10 minutes, with fire_probability (0 to 1, scale factor 0.01) and fire_result classes as described in Lesson 1; confirm field names against the current product user guide. Nigerian Navy raid reports are public news items; dates and LGAs are reliable, exact sites are often withheld. The Lillie Earth figures in this module (counts, thresholds, results) come from the company's September 2026 engineering log.

Previous10 Wildfire Systems End to EndNext11 MLOps and Cloud Deployment