MUTEX: Multimodal Urdu Toxic Span Detection
Fine-tuned XLM-RoBERTa with BIO token classification and CRF decoding, fused with MMS-300M
speech representations, over the URTOX corpus of 14,338 samples annotated to κ = 0.82.
Reaches 83.2% F1 on synthetic and 77.1% F1 on accent-stratified Urdu audio,
with token-level SHAP and attention heatmaps for explainability. Placed 6th at PIEAS Open Day 2026.
PythonPyTorchNLP
CRFMMS-300MExplainable AI
When Fusion Breaks: Robustness in Multimodal Emotion Recognition
Deleting audio and vision together cost nothing measurable, 96-101% of skill retained,
while removing text collapsed every model to 35-40%. The multimodal gain is largely illusory.
Six fusion architectures (early, late, TFN, LMF, MulT) measured across 14 corruption
operators and 32 evaluation axes on CMU-MOSEI's 22,856 utterances.
PyTorchCMU-MOSEIMultimodal Fusion
Robustness
Lacuņa: Research Gap & Discovery Engine
A deterministic engine over 16,605 NLP papers (73 languages, 26 tasks, 1,722 venues) that scores
under-researched language and task pairings out of 100, every figure traceable to the papers behind it.
Gazetteer tagging rather than model inference keeps counts reproducible. It surfaced that
48% of language-tagged papers never leave the highest-resource tier, and informed a
top 25 of 250+ finish at NeuroLogic '26 Global NLP Datathon.
NLPBibliometricsLow-Resource Languages
Vercel
Eternity Twin: Cognitive Interface for Model Interpretability
An interactive 3D interface for inspecting an emotion classifier's internal state, making visible the
gap between a model's inferred reasoning and its output behaviour, with uncertainty represented as a
first-class property of every node across 28 GoEmotions categories.
Engineered the systems layer for scale and access: Web Worker force-directed graph layout,
GPU-instanced whole-graph rendering in two draw calls, custom GLSL per-vertex region shading, and full
keyboard and screen-reader parity in WebGL (axe-clean, WCAG AA).
Next.jsReact Three FiberGLSL
TypeScriptWebGL
AEGIS: Agentic ISR Threat Intelligence Platform
A 7-agent defence intelligence system for UAV surveillance with telemetry classification, RAG-based
doctrine retrieval and SHAP explainability, reaching 0.89 F1 threat detection at ~1.8 s latency.
Shipped as a FastAPI and Streamlit dashboard with automated SALUTE reporting over 500+ simulated ISR
missions, achieving 0.84 Precision@3 retrieval and a sub-3% hallucination rate.
PythonFastAPILangGraph
SHAPChromaDBStreamlit
Multi-Agent RAG with Self-Correcting Pipelines
Nine specialised LangGraph agents handling query routing, retrieval and multi-step reasoning,
with hybrid ChromaDB and Tavily retrieval plus self-correction loops that cut hallucinations by 40%
against a baseline RAG pipeline.
LangGraphLangChainGroq
TavilyChromaDB
MorphoLens: Morphology in Multilingual Models
Do multilingual models actually understand morphology, or do they mostly memorise word forms? An
interactive research workspace built to interrogate that question across high and low-resource languages.
Runs on UniMorph data with a reproducible lemma-overlap experiment, spanning Turkish,
Urdu, Evenki, Chukchi and Romanian. Every surface form traces back to its source record.
UniMorphMorphologyLow-Resource NLP
Reproducible Experiments