Igor Joshevski

Igor Joshevski

Senior AI Integration Engineer · LLMs, RAG & Voice AI

Nuremberg, Germany

Sixteen years in production software. I ship AI into products that are already carrying real traffic—RAG search serving ~1M users/day, a real-time voice pipeline running at ~300 ms across 14 languages, and LLM-driven quality gates in CI that lifted throughput ~25% while cutting post-release defects ~60%. Python for LLM orchestration and data work, Node.js/TypeScript (NestJS) for the services around it.

RAG & Hybrid Retrieval LLM Agents & Tool Use ASR / TTS / VAD Python NestJS & TypeScript Docker & Kubernetes
16 yrs
in production software
~1M
users/day on RAG search
~300 ms
voice pipeline latency
80
developers previously managed

What I Work With

AI

AI & LLM

Retrieval systems that stay grounded, agents that use tools reliably, and speech pipelines fast enough to feel live. I work with Claude, Vertex AI, and self-hosted Llama via Ollama, and I fine-tune small models when a task deserves its own.

RAG & hybrid retrieval, pgvector, embeddings
LLM agents & tool use, prompt engineering
ASR/TTS, VAD, ONNX runtime optimisation
Fine-tuning · Claude, Vertex AI, Llama via Ollama
BE

Backend & Infrastructure

Python for LLM orchestration and data work, NestJS/Node.js for the real-time services around it. Everything ships through Docker and CI/CD, on CPU or CUDA, local or cloud— the voice work runs with no GPU dependency at all.

Python, TypeScript/Node.js, NestJS
REST & WebSocket APIs, microservices
PostgreSQL, BigQuery
Docker, Kubernetes, CI/CD, CPU & CUDA deployment

Angular/React where product work requires it—years of it behind me, but it is no longer the point.

How I Think About Technology

Philosophy and approach to software development and AI system design.

A demo is not a system. Most of my work is the distance between the two: what happens when a retrieval pipeline meets a 2M-product catalogue that changes daily, or when a speech model has to answer inside 300 milliseconds on a machine with no GPU. That distance is where the interesting engineering lives, and it is usually less about the model than about everything around it.

Grounding is the part I care about most. A RAG system is only as good as its weakest assumption about freshness, so I build the re-verification cycle before I build the clever ranking. The same instinct applies to LLMs inside a development process: putting Claude and self-hosted Ollama models into CI as quality gates only worked because the gates were measurable—throughput up ~25%, escaped defects down ~60% across 40 developers.

Sixteen years in, having managed 80 developers across 12 teams and been a founding engineer three times over, I've stopped believing that adoption is a tooling problem. It's a training and workflow problem wearing a tooling costume. Shipping the capability is half the job; the other half is the skills, agent tooling, and habits that make people actually reach for it.

AI & LLM Work

Systems I've taken from proof of concept to production, and what they had to survive to get there.

  • Retrieval at Scale
    In production

    Production RAG Platform

    2M-product catalogue · ~1M users/day

    Hybrid retrieval over a 2M-product catalogue: pgvector similarity search combined with BigQuery text search for multi-attribute matching across colour, size, features, and description.

    Owned embedding strategy, chunking, and context management, with a weekly re-verification cycle against source providers keeping answers grounded in current inventory.

    Keywords: RAG, Hybrid Retrieval, pgvector, BigQuery, Embeddings, Chunking
  • Voice AI
    Delivery Q4 2026

    Real-Time Speech Translation with VAD

    14 languages · ~300 ms processing latency

    Any-to-any real-time speech translation across 14 languages at ~300 ms processing latency (1.5 s to first audio including buffering), combining VAD-based segmentation with open-source ASR and ONNX-optimised translation/TTS.

    Validated on CPU-only and CUDA targets, local and cloud — no GPU dependency.

    Keywords: ASR, TTS, VAD, ONNX Runtime, Low Latency, CUDA & CPU
  • LLMs in the SDLC
    In production

    Enterprise AI Adoption

    40 developers · ~25% throughput · ~60% fewer escaped defects

    Integrated Claude and self-hosted Ollama models into CI for automated code review, critical-issue detection, and LLM-driven quality gates across a 40-developer organisation.

    Paired with skills training, agent tooling, and workflows — the adoption work that decides whether the tooling gets used at all.

    Keywords: Claude, Ollama, LLM in CI, Automated Code Review, Agent Tooling
  • Cost & Infrastructure
    In production

    LLM Routing & Cost Optimisation

    n8n automation platform · complexity-based routing

    Stood up a self-hosted n8n automation platform end to end — infrastructure, LDAP-backed authentication, and workflow design — then cut token spend by routing requests on complexity instead of defaulting everything to the most capable model.

    First month in production: ~30% of traffic classified low-complexity and served by self-hosted open-source models, 40–50% mid-complexity by mid-tier commercial models, and only ~20–30% reaching frontier models.

    Keywords: n8n, LLM Routing, Cost Optimisation, LDAP, Self-Hosted Models
    Read the write-up →
  • Model Work

    Model Fine-Tuning

    SmolLM v2 · task-specific behaviour

    Hands-on fine-tuning of small language models for task-specific behaviour, including dataset preparation and evaluation of output quality against base-model baselines.

    Small models earn their keep when the task is narrow and the latency budget is tight.

    Keywords: Fine-Tuning, SmolLM v2, Dataset Preparation, Evaluation
  • Voice AI
    Open source

    Meeting Transcriber

    Standalone CLI · Windows & macOS · fully local

    Real-time transcription of whatever's making noise on the machine — Teams, Zoom, or anything else — as a standalone CLI, not a plugin. Silero VAD segments speech, whisper.cpp transcribes it, all on-device.

    Capture and inference run on separate threads connected by a queue, so a slow transcription pass never drops or blocks audio capture.

    Keywords: Real-Time Transcription, whisper.cpp, Silero VAD, Local ASR, Open Source
    Read the write-up →

Track Record

Sixteen years of production software, most recently spent putting AI into products that already had users.

  1. Senior Full-Stack Developer — AI/ML Practitioner

    May 2025 – Present

    Avenga

    • Deliver customer-facing solutions across media, IoT, gaming, and telco, owning architecture from design through production rollout and working directly with client stakeholders.
    • Lead AI adoption across a 40-developer group and ship AI capability into client products, from proof of concept to production.
    • Build real-time APIs with NestJS/Node.js and WebSockets, performant Angular/TypeScript UIs, and Python services for data processing and LLM orchestration, shipped via Docker and CI/CD.

    Clients: IBM · Robert Bosch · Sunrise · Allwyn

  2. Senior Full-Stack Developer — Media, AI & iGaming

    Sep 2023 – May 2025

    Qinshift (now Avenga)

    • Delivered a Netflix-class streaming platform serving 5M users at launch, leading architecture and implementing micro-frontends over microservices on Kubernetes from concept to first production release.
    • Achieved sub-100 ms end-to-end latency for iGaming data ingestion at 3–4M concurrent users, designing a resilient multi-provider pipeline with optimised queuing, caching, and parallel processing.
    • Introduced the platform's first AI capability, taking RAG-backed product search from proof of concept to production.

    Clients: Deutsche Glasfaser · Cashpoint · Rumble · Lumis

  3. Senior Software Engineering Manager

    Aug 2017 – Sep 2023

    Seavus (now Avenga)

    • Managed 80 developers across 12 teams, improving delivery predictability and retention through structured capacity planning, career progression frameworks, and technical growth programmes.
    • Owned engineering quality standards and delivery procedures across the portfolio, standardising review, testing, and release practice.
    • Stayed hands-on across the PEAN stack (PostgreSQL, Express, Angular, Node.js) and React Native, with Docker and cloud deployment.

    Clients: Marginalen Bank · CSGi · IKEA · Hostopia

  4. Founding Engineer — Freelance

    2019 – 2024

    Built products for early-stage startups outside of full-time employment, owning the full path from idea to shipped solution.

    • AskBoss — React interface for their AI platform: visualisation of model behaviour plus the working surface for data classification, segmentation, and model testing.
    • TherapistHat — React Native mobile app for online training and continuing education of therapists in the US market.
    • Inky — React Native trivia mobile game, built end to end from concept to store release.

Earlier

  • Senior Web Developer → Technical Lead · Seavus · 2013 – 2017 — owned architecture for enterprise delivery teams and stepped up from senior engineer to technical lead within two years.
  • PHP Developer → Senior PHP Developer · Mikkelsen Media, Invideous Limited · 2009 – 2013 — architected a video streaming platform and player add-ons, and tuned backend systems for scalability under high request volume.
  • Technical Trainer · Seavus Education and Development Center · 2016 – 2020 (part-time) — trained and mentored developers in front-end technologies.

Education & Certifications

B.Sc. Information Technology — University "St. Kliment Ohridski", Bitola

AI Capabilities and Limitations · Database Developer · Apache 2.4 Administration · Linux Performance Tuning

Languages & Recognition

English (fluent) · Macedonian (native) · Serbian (fluent) · German (elementary)

Best Innovation Prize · Technical writing on Medium

Let's Connect

Always up for a conversation about retrieval that stays grounded, latency budgets in voice pipelines, or getting LLMs usefully into a team's workflow. Germany — reachable on LinkedIn.