Skip to content

Hossam
El-Kharbotly

Sovereign AIOn-premise infrastructureArabic NLP

I build AI agents, and I run the infrastructure they sit on.

AI Engineer with 3+ years taking Generative AI from prototype to production: multi-agent systems that take real actions, hybrid RAG over enterprise data, and the sovereign on-prem GPU stack that serves them to hundreds of concurrent users. Hands-on Arabic NLP.

Role
AI Engineer
Based in
Cairo, Egypt (currently on assignment in Amman, Jordan)
Status
Open to relocation & remote
Languages
Arabic (native) · English (fluent, IELTS 7 / C1)
01 / About

What I actually do

I work across both halves of production AI: multi-agent systems that take real actions for the people using them, and the on-premise infrastructure serving those systems. At Dar Al-Handasah I designed the model-agnostic on-premise LLM stack, presented the architecture to the CIO, and operate it: a fully in-house set of reasoning, embedding, reranking, OCR, and extraction models, so confidential project data never leaves the building and there is no per-token external API spend. I report to the CIO and a firm partner.

I also sit with the people who will use what I build. For Dar's PMC unit I worked on-site in Amman with the partner, the unit director, his directors, and the cost engineers to learn their use cases and business case, then demoed, tested, and trained the team on the agents face to face.

Most engineers are one or the other. Application engineers call somebody's API and cannot tell you why inference falls over at concurrency. Infrastructure engineers keep GPUs busy but have never shipped an agent a real user depends on. Owning the whole path is what lets me make the calls that only show up at the seams: how much context window a model can afford given the GPU and the real token profile, when to route a document region to a VLM instead of classical OCR, when an agent should stop and ask a human instead of guessing.

I work natively in Arabic, which matters in a region where Arabic-language AI is a first-class requirement rather than a localisation afterthought.

02 / ExperienceTap to expand

Where I've built it

Dar Al-Handasah

Feb 2026 to Present

Reasoning x2Qwen · vLLM
EmbeddingTEI
RerankTEI
OCRVLM
ExtractionQwen
GPU fleet3 x A100, on-prem
4 GPU serversA100 · RTX 4090
Hundredsconcurrent users at peak
Monthsuptime, no OOM
  • Embedded on-site with Dar's PMC unit in Amman: ran discovery with the PMC partner, the unit director, his directors, and the cost-engineering team to capture their use cases and business case, then demoed, tested, and onboarded the agents face to face, training ~10 cost engineers.
  • Build production multi-agent systems on AIDA, Dar's firm-wide AI platform, serving hundreds of users across 3 departments with thousands of requests per month.
  • Designed the model-agnostic on-premise LLM stack, presented the architecture to the CIO, and operate it: a hybrid (Azure + on-premise) estate of 6+ servers, including 4 GPU servers spanning A100 and RTX 4090 hardware, running open-weight LLMs (Qwen-family today) plus embedding, reranking, and specialty models on vLLM and TEI.
  • Built the BoQ Cost Agent used by cost engineers across PMC: a historic-cost lookup that took up to half a day now returns a benchmark, cross-project comparison, and fair-price estimate in minutes, with a citation trail the engineer can audit before signing off.
  • Tuned context windows, concurrency, and GPU memory to co-locate multiple models per GPU, serving hundreds of concurrent users and running months without OOM. Models are chosen through golden-set benchmarks against Azure OpenAI.
  • Integrated LiteLLM as the enterprise AI gateway: model routing, personalised API keys, and per-user budgets.
  • Built a human-in-the-loop layer that lets agents pause to ask the user for missing context, now a shared AIDA capability other teams build on.
  • Shipped the firm's first document-digitisation platform, turning a 3,500-page document into searchable data in about 25 minutes at ~100 files/day for 6 concurrent users.
  • Instrumented the AI estate with Langfuse tracing plus Prometheus, Grafana, and Loki, with auto-heal restarts. Review code for AIDA contributors.
  • Mentored ~30 trainees across 3 internship cohorts in Dar's Egypt, Lebanon, and Jordan offices.

Edulga

Aug 2025 to Jan 2026

Independent

Mar 2025 to Aug 2025

Al Farid Scan

Mar 2024 to Mar 2025

03 / ProjectsTap a card

Systems in production

04 / Stack

What I work with

Agents & LLM Systems

Multi-agent orchestrationLangGraphGraph engineeringAgent loopsHarness EngineeringHuman-in-the-loop pipelinesRAGGraphRAGText-to-SQLFine-tuning LLMs (LoRA/PEFT)DSPyGolden-set evaluation

Arabic AI & NLP

Arabic NLPArabic ASR / TTSArabic NERArabic semantic searchDialectal Arabic

Serving & Infrastructure

vLLMTEIGPU fleet managementMulti-model per GPUConcurrency tuningFastAPIDockerAzureAzure DevOpsAzure AI FoundryAWSCI/CD

Observability

LangfuseLiteLLM (model routing)MLflowPrometheusGrafanaLokiOpenTelemetryAuto-heal

Data & Vector Stores

Neo4jWeaviateQdrantPineconeMongoDBSQL ServerPostgreSQLFAISSMinIO

ML & Vision

PyTorchTensorFlowHuggingFace TransformersScikit-LearnOpenCVYOLO

Languages

PythonTypeScriptSQLCypherC++

Web & APIs

FastAPIReactNode.jsREST
06 / ContactOpen to relocation & remote

Let's talk

Hossam El-Kharbotly · AI Engineer · Cairo, Egypt (currently on assignment in Amman, Jordan)

Open to relocation, or remote from anywhere.