Explainable AI text classification
The prototype produces explainable AI text risk classifications rather than relying on a label alone.
Not sure where AI fits in your business? Book a full day on site AI Discovery visit.
Small language model case study
A fine tuned small language model for explainable AI text detection and adversarial robustness.
The rapid proliferation of AI generated text across professional and institutional environments has introduced a new class of operational risk. Organizations can no longer assume that written submissions, published content, or support communications are human authored and the consequences of misclassification vary significantly by context.
Project classification: This prototype explores explainable and adversarially tested AI text classification using a fine tuned small language model.

The prototype produces explainable AI text risk classifications rather than relying on a label alone.
Confidence outputs support risk based review instead of automatic conclusions.
Adversarial testing evaluates how the classifier behaves on paraphrased and rewritten AI text.
01 / Business context
The rapid proliferation of AI generated text across professional and institutional environments has introduced a new class of operational risk. Organizations can no longer assume that written submissions, published content, or support communications are human authored and the consequences of misclassification vary significantly by context. Hiring teams encounter AI generated job applications and assessments. Academic institutions handle AI assisted submissions.
02 / Challenge
Detection must remain accurate even after deliberate paraphrasing or humanization. This adversarial robustness requirement fundamentally reshapes training methodology, dataset design, and feature engineering.
03 / Workflow transformation
04 / Solution
The solution is built on a fine tuned RoBERTa classifier using LoRA (Low Rank Adaptation), trained on a hybrid dataset and enhanced with explainability and robustness layers. Model Selection: RoBERTa with LoRA RoBERTa was selected for its strong sequence classification capabilities and sensitivity to stylistic patterns distinguishing AI from human text.
Combine public, human written, AI generated and rewritten examples across domains.
Adapt an efficient classifier without full model retraining.
Return risk scores and interpretable text level signals instead of only a label.
Evaluate paraphrased and humanised text alongside false positive behaviour.
05 / Example workflow
06 / Delivery scope
The prototype produces probabilistic risk signals and must not be treated as proof of authorship or misconduct. No benchmark value is published without dataset and false positive documentation.
07 / Architecture and controls
Detection output is a risk signal and must not be treated as conclusive proof of authorship.
Human written and adversarially rewritten text must be included when evaluating model behaviour.
The interface exposes confidence and supporting indicators for reviewer interpretation.
Consequential decisions require independent evidence and qualified human review.
08 / Business value
The prototype produces explainable AI text risk classifications rather than relying on a label alone.
Confidence outputs support risk based review instead of automatic conclusions.
Adversarial testing evaluates how the classifier behaves on paraphrased and rewritten AI text.
Broader evaluation helps identify false positive risk on human written content.
A fine tuned smaller model demonstrates a path to lower latency and lower cost inference.
09 / Technology
Technology choices from the supplied project brief, mapped to the workflow each component supports.
Fine tuned text classification model
Reviewer or product user interface
AI, data processing and backend logic
Backend APIs and workflow orchestration
Structured schema and output validation
Model training and inference
Tabular data processing and analysis
Parameter efficient model fine tuning
Numerical processing
ON-SITE AI DISCOVERY
We come to your location for a full-day operational audit, then provide a prioritised AI Roadmap within 48 hours.