Senior ownership across the AI stack.
Agentic systems
Design and ship multi-agent workflows with explicit state, recovery, and evaluation — not freeform prompt chains.
Enterprise RAG
Build retrieval pipelines that survive real corpora — chunking, embeddings, hybrid search, evaluation, and answer grounding.
LLM evaluation
End-to-end suites with LLM-as-judge, regression gates, and human review — trust comes from measurement.
Arabic-first AI
Production Arabic NLP: tokenization, prompt structuring, evaluation, and content generation alongside English.
Production delivery
Python, FastAPI, Azure, Postgres, Playwright. Deploy, observe, debug, and harden under real traffic.
Production systems, with the receipts.
Sanitized case study
Role · AI-layer lead — architecture and implementation
Production multi-agent system generating Arabic-first content for government workflows, with deterministic orchestration and an evaluation suite gating every release.
PythonAutoGenAzure OpenAIWebSocketsPlaywrightPostgreSQL
Sanitized case study
Role · Designed and evaluated end-to-end
A multi-stage destination-discovery and award-feasibility engine for a Gulf airline's loyalty programme — turning a member's intent into ranked, explainable destination and routing suggestions, benchmarked against a baseline before shipping.
PythonAzure OpenAIHybrid RAGBatched availability APIEvaluation harness
Sanitized case study
Role · AI backend lead — architecture and implementation
Production multi-agent system that ingests inbound communications requests, classifies and validates them against policy, and routes each to the right downstream workflow — behind a strict state machine and data-residency guarantees.
Python 3.12FastAPIMicrosoft Agent FrameworkAzure OpenAIPydantic v2pytest
Sanitized case study
Role · POC lead and integration spec author
A generation pipeline that turns a brief into short cinematic video with native Arabic narration and on-screen subtitles — combining video, image, text, and speech models behind one orchestrated workflow.
Azure OpenAIVeo (video)Gemini (image)Azure AI SpeechPythonWebSockets
Case study
Role · Implemented end-to-end
A retrieval-augmented assistant grounded in Saudi legal documents from the Ministry of Justice, SDAIA, and municipalities — delivered end-to-end as a public-facing assistant.
PythonLangChainAzure OpenAIFAISSMultilingual embeddingsCross-encoder re-ranker
Shipped roles, in order of recency.
Ghaia.ai
Senior AI Engineer
Oct 2025 — Present
Remote · Doha, Qatar
Architect and ship production multi-agent systems for public-sector and enterprise clients in Arabic and English.
- ArchitectedArchitected the AI layer of a WebSocket-driven AutoGen system generating Arabic-first government content.
- ImplementedImplemented a declarative, state-gated prompt-composition engine, replacing a large legacy prompt builder with independently testable sections.
- OwnedOwned an agentic end-to-end suite with LLM-as-judge scoring for Arabic outputs.
- ContributedContributed evaluation tooling, retrieval improvements, and production debugging across public-sector and enterprise workflows in the Gulf.
Build custom AI systems for SMB and public-sector clients — RAG, automation, and applied LLM pipelines.
- ImplementedImplemented Al Qanoon — an Arabic-first RAG assistant grounded in publicly available Saudi legal texts.
- LedLed delivery of bespoke AI workflows including outreach automation and content generation pipelines.
- EvaluatedEvaluated retrieval quality and grounding across Arabic and English corpora.
Earlier experience
- Codingal — Senior Coding InstructorJul 2022 — Feb 2025 · Remote
- Ribat Store — Head of ITSep 2022 — Nov 2023 · Pakistan
- Adrenod — Chief Technical OfficerMay 2023 — Aug 2023 · Islamabad, Pakistan
- Webspire — Co-FounderMay 2021 — May 2022 · Remote
The stack behind the work.
Agentic systems
- AutoGen
- Microsoft Agent Framework
- LangGraph
- State machines
- Tool orchestration
- Recovery & retries
Retrieval and RAG
- FAISS
- Vector search
- Hybrid retrieval
- Chunking strategies
- Re-ranking
- Citations & grounding
LLM evaluation
- LLM-as-judge
- E2E test suites
- Regression gates
- Prompt versioning
- Human-in-the-loop review
Backend & infrastructure
- Python
- FastAPI / Flask
- PostgreSQL
- SQLAlchemy
- WebSockets
- Azure OpenAI
- Microservices
Applied AI
- Arabic NLP
- Document understanding
- Multimodal generation
- Workflow automation
- Public-sector deployments