SONAR

Methodology

The honest spine

Every number SONAR reports traces back to the pool that produced it. This page states how the pool is built, how it is weighted, and where the method refuses to speak.

Dome confirmation

The dome is the set of accounts confirmed to belong to the place being measured. An account enters it on resident evidence, not on a keyword or a location tag.

The confirmed-pool size and the decided-pool size are reported on every output. A small dome is stated as a small dome.

Impact weighting and the five percent cap

Standing is not a count of posts. Each confirmed account contributes by realized engagement first and reach second, so one loud account does not outvote a quiet neighborhood by posting more.

To keep any single account from dominating a thin pool, an anti-concentration cap holds each account to at most five percent of the pool's weight. Below the account count where the cap can bind, it cannot engage, and the output says so with a cap_infeasible flag rather than hiding the concentration.

Pool gates and what inconclusive means

A headline is published only when the pool can support it. Two floors apply: a minimum count of decided signals, and a minimum count of decided accounts, so a handful of accounts posting many times does not clear the gate on volume alone.

When a floor is not met, the standing reads inconclusive or pool building, with the reasons attached: thin, concentrated, cap infeasible, or no calibration record. Inconclusive is a real answer. It is reported plainly, never dressed as a number.

Inauthenticity screening

Before weighting, the pool is screened for coordinated and inauthentic behavior: posting bursts, duplicated text, and freshly created accounts. The screen flags, it does not silently remove.

Both the unfiltered standing and the standing with flagged accounts removed are shown together, so the effect of the screen is visible rather than assumed. Volume from a coordinated cluster is reported as coordinated, not as breakthrough.

Stance classification and accuracy checks

Each post about a candidate is classified as supportive, opposed, or neutral. The classifier reads any language and returns a stance in English. Each market carries a short language context so local idiom is read correctly rather than dropped as noise. In Gambia, for example, the Wolof phrase for asking an incumbent to step down is read as opposition, not as support.

Classifier accuracy is checked against a labeled review and published with the source of the labels stated distinctly:

Alongside agreement, an attribution error rate is reported: the share of reviewed posts that were wrongly attributed to a candidate in the first place. A misattribution is an error to publish, not to bury.

What SONAR refuses to do

No demographic modeling. Slices are by platform, geography, and community source only. SONAR does not infer age, race, or gender and does not weight by them.

No turnout claims. SONAR measures conversation among confirmed residents. It does not project who will vote.

No verdict without a calibration record. In a market with no scored calibration pair, the standing renders inconclusive by rule. Confidence comes from pool volume and the published record, never from model features or narrative fit.

How a lock is scored

A prediction becomes a lock the moment it is written, and a lock is write-once. The candidate shares, the methodology version, the data window and the election date are fixed in one row that cannot be edited afterwards. The row's hash is published, and the hash is stamped into the Bitcoin blockchain, so the claim that a reading existed before a result is checkable by anyone without trusting SONAR.

Scoring happens after the official result is certified, never before, and it uses the certified figures rather than a media call. Four numbers are recorded for every scored pair: whether the winner was called correctly, the mean absolute error against the certified shares, the signed error, and the margin error. Where a market has a polling benchmark-of-record fixed at lock time, the same comparison is run against the polls, so the comparison against the polls is answered with a number rather than an adjective. A market with no benchmark is reported as having none.

A scored pair enters the public ledger whatever it says. A missed call stays on the record beside the calls that landed.

What SONAR measures, and what a poll measures

A poll asks a sample of people a question and reports what they answered. SONAR reads what residents said on their own, unprompted, and reports how that conversation is distributed. These are different measurements, and neither is a substitute for the other. A poll reaches people who do not post; SONAR reaches people who would never answer a poll. A poll's error is dominated by who agrees to respond and how the question is worded; SONAR's is dominated by who chooses to speak and how loudly.

Because the two instruments fail in different directions, the interesting question is not which is right but where the result lands relative to both. That is what the graphic below shows: the benchmark poll fixed at lock time, the raw conversation reading, and — once the result certifies — the certified value on the same share line.

The benchmark of record is fixed when the prediction locks and never changed afterwards, so the comparison cannot be re-chosen once the result is known.

Ground truth and independent annotation

Independent human annotation: in procurement. Classifier accuracy is currently checked two ways: against a labelled review by the operator, and against a second, larger model reading the same sample. The second is reported as a cross-model check and is never described as human validation. An independent annotation vendor — labels produced by people with no stake in the result — is being procured and has not yet delivered a batch. Until it does, no figure on this site is described as independently validated.

Maturity: what is proven, and what is not

Every claim SONAR makes carries a maturity stamp, because a method tested twice is not a method tested twenty times.

Measured. The pool construction, weighting, gates and screens described above run in production across every deployed market. These are engineering, and they do what this page says they do.

Calibrated on a small number of pairs. The scored record is short. Each new scored pair can move the error estimate materially, and the published ledger shows exactly how many pairs stand behind any accuracy figure.

Hypothesis. The loudness correction below, and any claim that a correction fitted in one country transfers to another, are hypotheses under test. They are labelled as hypotheses wherever they appear and they never silently become the headline.

Limitations, stated plainly

This is an online electorate, not the electorate. SONAR measures residents who post about a race on public platforms. That population is younger, more urban and more engaged than the voting population almost everywhere, and in some markets it sits overwhelmingly on a single platform. Every output reports the confirmed and decided pool sizes so the reader can see how narrow the base is. A standing is a measurement of the conversation, and the conversation is not the ballot box.

Loud voices skew the reading, and the correction is itself under test. The people who post most are not representative of the people who agree with them. Impact weighting and the five percent cap blunt this, but they do not remove it: a leader's raw share typically reads hotter than the certified result. SONAR estimates that skew from its own scored pairs and publishes both the raw reading and the corrected one, never the corrected one alone. The correction is fitted on a handful of pairs from a small number of races. Whether it holds in a new country is an open question, and the first such test is reported as a test whichever way it lands.

Silence is not measured, and abstention is a real answer. A resident who never posts is invisible here, and a decided pool of ninety accounts is ninety accounts, not a sample of a nation. When the pool cannot carry a question, SONAR abstains: it returns inconclusive, or renders no figure on a thin cell, rather than producing a number with the precision of a real one. Abstention is applied by fixed floors set in advance, so it is never chosen after seeing whether the answer was convenient.

Check it yourself

None of this needs to be taken on trust. The locked predictions, their hashes, their blockchain stamps and the scored results are published, and the verification procedure runs with standard tools and no account. See verify the record for the exact steps, and the calibration record for every pair scored so far, including the ones that missed.

The record that backs all of this is public. See the calibration record.