This site is still being built — some things may change or break.

Local AI misinformation classification system.

My bachelor's thesis at the University of Lübeck.

Development of a browser extension for
to display warnings about misinformation

Development of a browser extension for X to display warnings about misinformation

Bachelor Thesis

My bachelor's thesis at the University of Lübeck.

Fact-checking on social media is slow and selective — a few posts get a note, the rest scroll by unchecked. This thesis explores the opposite: a clear, explainable credibility signal on every post. The result is a browser extension for X that classifies posts in real time with a large language model running entirely on your own device, so nothing ever leaves your browser. Each post is sorted into one of five categories and explained through layered, tap-to-reveal detail. Built with a human-centered design process and tested head-to-head against X's own Community Notes, it was judged more transparent, more systematic, and harder to manipulate.

Highlights

A MacBook showing the X feed with the extension's classification label and explanation panel opened on a post.
Automated content classification for your entire feed
A screenshot of a web-based analytics dashboard on a MacBook. The 'X Misinformation Warning Prototype' shows a summary of 817 analyzed posts, a line graph of classification trends over time, and a gallery of recently analyzed social media posts with credibility tags.
Dashboard providing insights into engagement with misinformation
Offline by design. Privacy by default.
User interface elements for a misinformation detection tool. It includes color-coded status labels, a short summary warning box , and a detailed Classification window providing credibility scores and analysis.
Designed for clarity first, the interface lets you gradually reveal more detailed explanations - so you can understand decisions at your own pace and depth

Classification
Five categories.
One clear signal.

A laptop on a table showing a social media feed with automated misinformation labels applied to various posts.

Five verdicts, not two. Real claims rarely split cleanly into true and false. Every post is sorted into one of five categories — True, False, Disputed, Unclear or Parody — so nuance is never forced into a binary.

A scale, not a switch. The five categories sit on a single spectrum, each carrying its own colour and icon, so a verdict registers at a glance — before you read a single word.

On-device by design. Classification runs on a language model built into the browser, on your own machine — post text, images and prompts never leave it. Declared ads and parody accounts are filtered out before the model even runs.

Tap to go deeper. Progressive disclosure keeps the feed calm: a compact badge first, a one-line reason on tap, and a full detail card with analysis, red flags and sources only when you want it.

A false-flagged post's detail view showing 0-to-10 bars for truthfulness and harm potential.

Measured, not just labelled. When a post is flagged false, it is scored from 0 to 10 for both truthfulness and harm potential — turning a single verdict into something you can actually weigh. Other categories stay deliberately unscored.

An author profile with a credibility score aggregated from their recently classified posts.

Credibility, earned over time. From at least three classified posts, an author's recent history is aggregated into a credibility signal, weighted toward newer activity — context that no single tweet could ever give.

A post whose image quietly contradicts its text, with the model's image description shown beside the verdict.

It reads the picture, too. When a post carries an image, the model describes it and weighs it against the words — catching the cases where a visual quietly contradicts the claim.

Built for every kind of sight. High-contrast, colour-blind and icon-only modes keep every warning legible — and no verdict ever leans on colour alone to make its point.

A minimalist workstation with a Mac Pro tower, a large central monitor, and a MacBook Pro. The central screen shows a web browser with the extension flagging a post in an X feed as misinformation, with scores for truthfulness and credibility. The laptop displays the extension's analytics dashboard tracking hundreds of analyzed posts with their classification categories.

The Research
A human-centered process.
Ten users put it to the test.

A circular human-centered design process diagram with four stages feeding back into one another.

A method, not a hunch. The whole project followed the ISO 9241-210 process for human-centered design — analyse, conceive, build, evaluate, then loop back — so every decision traces to a user need rather than a guess.

A funnel narrowing from hundreds of database search results down to a handful of core sources.

Grounded in the literature. Three systematic searches across Google Scholar, the ACM Digital Library and Scopus narrowed hundreds of hits to a handful of core sources — setting how the warnings should look and how they should explain themselves.

Annotated interface sketches showing how an explanation expands across three levels of detail.

Explainable by construction. The interface is built on established explainable-AI principles — naturalness, responsiveness, flexibility and sensitivity — with progressive disclosure as the organising idea, so explanation scales with the reader's attention.

A chart contrasting stable scores on false posts with widely scattered scores on neutral posts.

A score I could defend. A functional test exposed the numeric score: reliable once a post was already judged false, but scattered on neutral content. So scoring was deliberately limited to false posts — a constraint, made honest.

Before-and-after Figma prototype screens carrying annotations from a formative evaluation.

Tested early, changed often. A formative evaluation of the Figma prototype reshaped the system: two categories were added, high-contrast and colour-blind palettes introduced, and the dashboard rebuilt around clearer explanations.

A participant in a usability lab using the extension while thinking aloud, observed through a one-way setup.

Studied with real users. A summative, mixed-methods study (n = 10, within-subjects) in the usability lab paired think-aloud sessions and interviews with a multidimensional trust questionnaire — and put the extension head-to-head with X's Community Notes.

A bar chart of trust scores across the questionnaire's dimensions, all sitting in the upper range.

The verdict on trust. Overall trust landed at 3.44 out of 4, strongest on usefulness and intention to use. What participants described as trustworthy was consistent throughout: every post judged by one visible, unchanging logic.

A side-by-side comparison of the AI warning and a Community Note across trust dimensions.

More trusted than Community Notes. Against Community Notes, the extension was judged more transparent, more systematic and less open to manipulation — their selectiveness bred doubt. With only ten participants, the findings point a direction rather than settle it.

A roadmap sketch of planned features beyond the current build.

What's next. The thesis maps the road ahead: a prompt and feedback log for full auditability, an optional choice between local and cloud models, and explanations that adapt to each reader.