Hossam
El-Kharbotly
Sovereign AIOn-premise infrastructureArabic NLP
I build AI agents, and I run the infrastructure they sit on.
AI Engineer with 3+ years taking Generative AI from prototype to production: multi-agent systems that take real actions, hybrid RAG over enterprise data, and the sovereign on-prem GPU stack that serves them to hundreds of concurrent users. Hands-on Arabic NLP.
- Role
- AI Engineer
- Based in
- Cairo, Egypt (currently on assignment in Amman, Jordan)
- Status
- Open to relocation & remote
- Languages
- Arabic (native) · English (fluent, IELTS 7 / C1)
What I actually do
I work across both halves of production AI: multi-agent systems that take real actions for the people using them, and the on-premise infrastructure serving those systems. At Dar Al-Handasah I designed the model-agnostic on-premise LLM stack, presented the architecture to the CIO, and operate it: a fully in-house set of reasoning, embedding, reranking, OCR, and extraction models, so confidential project data never leaves the building and there is no per-token external API spend. I report to the CIO and a firm partner.
I also sit with the people who will use what I build. For Dar's PMC unit I worked on-site in Amman with the partner, the unit director, his directors, and the cost engineers to learn their use cases and business case, then demoed, tested, and trained the team on the agents face to face.
Most engineers are one or the other. Application engineers call somebody's API and cannot tell you why inference falls over at concurrency. Infrastructure engineers keep GPUs busy but have never shipped an agent a real user depends on. Owning the whole path is what lets me make the calls that only show up at the seams: how much context window a model can afford given the GPU and the real token profile, when to route a document region to a VLM instead of classical OCR, when an agent should stop and ask a human instead of guessing.
I work natively in Arabic, which matters in a region where Arabic-language AI is a first-class requirement rather than a localisation afterthought.
Where I've built it
Dar Al-Handasah
Feb 2026 to Present- Embedded on-site with Dar's PMC unit in Amman: ran discovery with the PMC partner, the unit director, his directors, and the cost-engineering team to capture their use cases and business case, then demoed, tested, and onboarded the agents face to face, training ~10 cost engineers.
- Build production multi-agent systems on AIDA, Dar's firm-wide AI platform, serving hundreds of users across 3 departments with thousands of requests per month.
- Designed the model-agnostic on-premise LLM stack, presented the architecture to the CIO, and operate it: a hybrid (Azure + on-premise) estate of 6+ servers, including 4 GPU servers spanning A100 and RTX 4090 hardware, running open-weight LLMs (Qwen-family today) plus embedding, reranking, and specialty models on vLLM and TEI.
- Built the BoQ Cost Agent used by cost engineers across PMC: a historic-cost lookup that took up to half a day now returns a benchmark, cross-project comparison, and fair-price estimate in minutes, with a citation trail the engineer can audit before signing off.
- Tuned context windows, concurrency, and GPU memory to co-locate multiple models per GPU, serving hundreds of concurrent users and running months without OOM. Models are chosen through golden-set benchmarks against Azure OpenAI.
- Integrated LiteLLM as the enterprise AI gateway: model routing, personalised API keys, and per-user budgets.
- Built a human-in-the-loop layer that lets agents pause to ask the user for missing context, now a shared AIDA capability other teams build on.
- Shipped the firm's first document-digitisation platform, turning a 3,500-page document into searchable data in about 25 minutes at ~100 files/day for 6 concurrent users.
- Instrumented the AI estate with Langfuse tracing plus Prometheus, Grafana, and Loki, with auto-heal restarts. Review code for AIDA contributors.
- Mentored ~30 trainees across 3 internship cohorts in Dar's Egypt, Lebanon, and Jordan offices.
Edulga
Aug 2025 to Jan 2026Independent
Mar 2025 to Aug 2025Al Farid Scan
Mar 2024 to Mar 2025Systems in production
What I work with
Agents & LLM Systems
Arabic AI & NLP
Serving & Infrastructure
Observability
Data & Vector Stores
ML & Vision
Languages
Web & APIs
Let's talk
The fastest way to reach me is WhatsApp or email. I reply the same day.