Igor Joshevski
Senior AI Integration Engineer · LLMs, RAG & Voice AI
Nuremberg, Germany
Sixteen years in production software. I ship AI into products that are already carrying real traffic—RAG search serving ~1M users/day, a real-time voice pipeline running at ~300 ms across 14 languages, and LLM-driven quality gates in CI that lifted throughput ~25% while cutting post-release defects ~60%. Python for LLM orchestration and data work, Node.js/TypeScript (NestJS) for the services around it.
What I Work With
AI & LLM
Retrieval systems that stay grounded, agents that use tools reliably, and speech pipelines fast enough to feel live. I work with Claude, Vertex AI, and self-hosted Llama via Ollama, and I fine-tune small models when a task deserves its own.
Backend & Infrastructure
Python for LLM orchestration and data work, NestJS/Node.js for the real-time services around it. Everything ships through Docker and CI/CD, on CPU or CUDA, local or cloud— the voice work runs with no GPU dependency at all.
Angular/React where product work requires it—years of it behind me, but it is no longer the point.
How I Think About Technology
Philosophy and approach to software development and AI system design.
A demo is not a system. Most of my work is the distance between the two: what happens when a retrieval pipeline meets a 2M-product catalogue that changes daily, or when a speech model has to answer inside 300 milliseconds on a machine with no GPU. That distance is where the interesting engineering lives, and it is usually less about the model than about everything around it.
Grounding is the part I care about most. A RAG system is only as good as its weakest assumption about freshness, so I build the re-verification cycle before I build the clever ranking. The same instinct applies to LLMs inside a development process: putting Claude and self-hosted Ollama models into CI as quality gates only worked because the gates were measurable—throughput up ~25%, escaped defects down ~60% across 40 developers.
Sixteen years in, having managed 80 developers across 12 teams and been a founding engineer three times over, I've stopped believing that adoption is a tooling problem. It's a training and workflow problem wearing a tooling costume. Shipping the capability is half the job; the other half is the skills, agent tooling, and habits that make people actually reach for it.
AI & LLM Work
Systems I've taken from proof of concept to production, and what they had to survive to get there.
-
Retrieval at ScaleIn production
Production RAG Platform
2M-product catalogue · ~1M users/day
Hybrid retrieval over a 2M-product catalogue: pgvector similarity search combined with BigQuery text search for multi-attribute matching across colour, size, features, and description.
Owned embedding strategy, chunking, and context management, with a weekly re-verification cycle against source providers keeping answers grounded in current inventory.
Keywords: RAG, Hybrid Retrieval, pgvector, BigQuery, Embeddings, Chunking -
Voice AIDelivery Q4 2026
Real-Time Speech Translation with VAD
14 languages · ~300 ms processing latency
Any-to-any real-time speech translation across 14 languages at ~300 ms processing latency (1.5 s to first audio including buffering), combining VAD-based segmentation with open-source ASR and ONNX-optimised translation/TTS.
Validated on CPU-only and CUDA targets, local and cloud — no GPU dependency.
Keywords: ASR, TTS, VAD, ONNX Runtime, Low Latency, CUDA & CPU -
LLMs in the SDLCIn production
Enterprise AI Adoption
40 developers · ~25% throughput · ~60% fewer escaped defects
Integrated Claude and self-hosted Ollama models into CI for automated code review, critical-issue detection, and LLM-driven quality gates across a 40-developer organisation.
Paired with skills training, agent tooling, and workflows — the adoption work that decides whether the tooling gets used at all.
Keywords: Claude, Ollama, LLM in CI, Automated Code Review, Agent Tooling -
Cost & InfrastructureIn production
LLM Routing & Cost Optimisation
n8n automation platform · complexity-based routing
Stood up a self-hosted n8n automation platform end to end — infrastructure, LDAP-backed authentication, and workflow design — then cut token spend by routing requests on complexity instead of defaulting everything to the most capable model.
First month in production: ~30% of traffic classified low-complexity and served by self-hosted open-source models, 40–50% mid-complexity by mid-tier commercial models, and only ~20–30% reaching frontier models.
Keywords: n8n, LLM Routing, Cost Optimisation, LDAP, Self-Hosted ModelsRead the write-up → -
Model Work
Model Fine-Tuning
SmolLM v2 · task-specific behaviour
Hands-on fine-tuning of small language models for task-specific behaviour, including dataset preparation and evaluation of output quality against base-model baselines.
Small models earn their keep when the task is narrow and the latency budget is tight.
Keywords: Fine-Tuning, SmolLM v2, Dataset Preparation, Evaluation -
Voice AIOpen source
Meeting Transcriber
Standalone CLI · Windows & macOS · fully local
Real-time transcription of whatever's making noise on the machine — Teams, Zoom, or anything else — as a standalone CLI, not a plugin. Silero VAD segments speech, whisper.cpp transcribes it, all on-device.
Capture and inference run on separate threads connected by a queue, so a slow transcription pass never drops or blocks audio capture.
Keywords: Real-Time Transcription, whisper.cpp, Silero VAD, Local ASR, Open SourceRead the write-up →
Technical Expertise Summary
Core competencies: RAG & hybrid retrieval, vector search & embeddings, pgvector, LLM agents & tool use, prompt engineering, fine-tuning, ONNX, ASR/TTS, VAD, Python, TypeScript/NestJS, PostgreSQL, BigQuery, Docker, Kubernetes
Domain experience: Media & streaming, iGaming, telco, IoT, e-commerce search, real-time platforms
Track Record
Sixteen years of production software, most recently spent putting AI into products that already had users.
-
Senior Full-Stack Developer — AI/ML Practitioner
May 2025 – PresentAvenga
- Deliver customer-facing solutions across media, IoT, gaming, and telco, owning architecture from design through production rollout and working directly with client stakeholders.
- Lead AI adoption across a 40-developer group and ship AI capability into client products, from proof of concept to production.
- Build real-time APIs with NestJS/Node.js and WebSockets, performant Angular/TypeScript UIs, and Python services for data processing and LLM orchestration, shipped via Docker and CI/CD.
Clients: IBM · Robert Bosch · Sunrise · Allwyn
-
Senior Full-Stack Developer — Media, AI & iGaming
Sep 2023 – May 2025Qinshift (now Avenga)
- Delivered a Netflix-class streaming platform serving 5M users at launch, leading architecture and implementing micro-frontends over microservices on Kubernetes from concept to first production release.
- Achieved sub-100 ms end-to-end latency for iGaming data ingestion at 3–4M concurrent users, designing a resilient multi-provider pipeline with optimised queuing, caching, and parallel processing.
- Introduced the platform's first AI capability, taking RAG-backed product search from proof of concept to production.
Clients: Deutsche Glasfaser · Cashpoint · Rumble · Lumis
-
Senior Software Engineering Manager
Aug 2017 – Sep 2023Seavus (now Avenga)
- Managed 80 developers across 12 teams, improving delivery predictability and retention through structured capacity planning, career progression frameworks, and technical growth programmes.
- Owned engineering quality standards and delivery procedures across the portfolio, standardising review, testing, and release practice.
- Stayed hands-on across the PEAN stack (PostgreSQL, Express, Angular, Node.js) and React Native, with Docker and cloud deployment.
Clients: Marginalen Bank · CSGi · IKEA · Hostopia
-
Founding Engineer — Freelance
2019 – 2024Built products for early-stage startups outside of full-time employment, owning the full path from idea to shipped solution.
- AskBoss — React interface for their AI platform: visualisation of model behaviour plus the working surface for data classification, segmentation, and model testing.
- TherapistHat — React Native mobile app for online training and continuing education of therapists in the US market.
- Inky — React Native trivia mobile game, built end to end from concept to store release.
Earlier
- Senior Web Developer → Technical Lead · Seavus · 2013 – 2017 — owned architecture for enterprise delivery teams and stepped up from senior engineer to technical lead within two years.
- PHP Developer → Senior PHP Developer · Mikkelsen Media, Invideous Limited · 2009 – 2013 — architected a video streaming platform and player add-ons, and tuned backend systems for scalability under high request volume.
- Technical Trainer · Seavus Education and Development Center · 2016 – 2020 (part-time) — trained and mentored developers in front-end technologies.
Education & Certifications
B.Sc. Information Technology — University "St. Kliment Ohridski", Bitola
AI Capabilities and Limitations · Database Developer · Apache 2.4 Administration · Linux Performance Tuning
Languages & Recognition
English (fluent) · Macedonian (native) · Serbian (fluent) · German (elementary)
Best Innovation Prize · Technical writing on Medium
Let's Connect
Always up for a conversation about retrieval that stays grounded, latency budgets in voice pipelines, or getting LLMs usefully into a team's workflow. Germany — reachable on LinkedIn.