George (Gosha) MareevAI Engineer · GMT+3
AI engineer · RAG · on-prem LLM infrastructure

LLMs, ready for production.

I build the platform layer behind enterprise AI — model gateways, RAG agents, secure deployment, evaluation, observability and prompt-level audit.

AI/ML engineer on an enterprise platform team · since June 2025

Stack Python · FastAPI · Docker · Postgres/pgvector · LiteLLM · Open WebUI · Loki · nginx · Linux

Model gateway. RAG agents. Observability. Audit.
Built for data that cannot leave the environment.
One production platform. Three engineering layers.
Scroll into the system

Flagship system · production

Enterprise on-prem
AI platform.

Designed and delivered end to end for an enterprise customer: the gateway, two RAG assistants, secure access, evaluation, operational telemetry and a separate prompt-audit path. Teams can use approved models and cited internal knowledge inside the secured environment, with requests traceable through a dedicated audit stream.

Client identity and source code remain private. The architecture, boundaries and engineering decisions below are safe to discuss in detail.

120enterprise users
10teams served
02production RAG assistants
SSOMicrosoft Entra ID at ingress
Architecture · simplified
Enterprise users
Entra ID + nginx
Open WebUILiteLLM gateway
On-prem / compatible models
RAG assistants + SharePointOperational logs → LokiPrompt audit → SIEM

Platform & secure deployment

A controlled path from users to models.

A self-hosted platform for an enterprise environment where sensitive data could not be sent directly to public model APIs.

What I built

  • LiteLLM as the single model gateway and Open WebUI as the user surface
  • nginx with TLS termination and SSO through Microsoft Entra ID
  • Four services operated independently, with only ingress exposed publicly
  • Separate databases and database users per service
  • Documented proxy handling for both the container runtime and image registry

Design decision

Identity, routing and policy stay at the platform boundary instead of being reimplemented inside each AI application.

Proof

Designed and delivered end to end; running in production with deployment runbooks, systemd units, GitLab CI and an operations cheatsheet.

LiteLLMOpen WebUIPostgreSQLnginxDocker ComposeEntra IDoauth2-proxy

Production RAG assistants

Retrieval designed to know its limits.

Two domain assistants share one core while preserving their own retrieval scope, content rules and user-facing behavior.

What I built

  • Metadata-aware retrieval across document category, system, audience and topic tags
  • Glossary-driven query expansion plus keyword overlap on top of vector scores
  • Reranking with a fallback when reranking degrades the selected context
  • Confidence scoring, grounded refusal and clickable source attribution
  • Incremental and full rebuild modes with manifest tracking and PDF fallback
  • SharePoint synchronization through Microsoft Graph with versioning

Design decision

A confident refusal is a product feature. When retrieval cannot support an answer, the assistant should expose that boundary rather than improvise.

Proof

A/B comparison captures latency, confidence and source counts, followed by expert QA review and feedback that can be fed back into ranking.

PythonChromaDBrerankersOllamaHuggingFaceMicrosoft Graphk6

Observability & prompt audit

Operational telemetry without leaking prompts.

Runtime health and security audit data take separate routes so prompt content never lands in the operational log store by accident.

What I built

  • Fluent Bit shipping container logs from journald to Grafana Loki over RFC5424
  • A separate prompt-level audit stream to the enterprise SIEM
  • Credential masking in Lua before operational logs leave the host
  • Filesystem-backed queues and disk-resident emitter buffers
  • An unprivileged container with all capabilities dropped and a read-only root filesystem
  • Prometheus metrics plus a SIEM audit field reference

Design decision

The safest sensitive log is the one that never enters the general logging pipeline. Separation is enforced by architecture, not only by redaction.

Proof

The audit transport and field contract are shipped. Detection rules for injection verdicts, policy violations and delivery control remain an explicit roadmap item.

Fluent BitGrafana LokijournaldRFC5424LuaPrometheusSIEM

Public engineering evidence

Explore the reference implementation.

A runnable, synthetic-data reference: cited RAG answers, scoped access, privacy checks and correlated audit events through Open WebUI and a single LiteLLM gateway.

Core RAG checks · each of 3 runs
32/32
Expanded RAG checks
125/128
Cloud.ru PII / secret values masked
35/50
Live privacy callback checks
7/7

Fixed synthetic tests, separate from production results; optional classifiers run in observation mode. See the reports for methodology, measured misses and deployment limits.

Engineering details · Production / Reference / Roadmap
Production experience, public reference and next steps have distinct scopes.
LayerProductionPublic referenceReference roadmap
Access & deploymentEntra SSO, isolated services and an on-prem model gatewayGoogle SSO, Open WebUI groups, two document scopes and one LiteLLM gatewayLive Entra verification, HTTPS deployment and deployment-specific policy mapping
RetrievalTwo domain assistants, semantic retrieval and SharePoint ingestionNative Knowledge RAG and pgvector; core 32/32 in three runs, expanded 125/128; corpus upgrade/rollbackRepresentative domain data, independent human review and stronger language fidelity
GuardrailsPlatform guardrail policy remains in progressEN/RU Presidio, removal of injected retrieved chunks, and observation-only Jev / Cloud.ru privacy checksBroader adversarial coverage and calibrated classifier enforcement
Media privacyDeployment-specific media policy requires separate validationLocal OCR/STT, sensitive-image blocking, sanitized speech text and a fixed 64-case EN/RU media evaluationBroader language coverage, visual personal data and representative OCR/STT evaluation
Telemetry & auditOperational logs to Loki and a dedicated prompt-audit transport to SIEMDurable bounded audit spool, private mTLS delivery, persistent checkpoints and event deduplicationEnterprise SIEM integration and resistance to privileged history rewriting

Context quarantine removes retrieved document chunks containing detected injected instructions before they reach the model. Jev is an optional semantic risk classifier evaluated on 128 cases; observation (shadow) mode records verdicts without using them to block requests or grant access.

Recorded local results · synthetic data · evidence replay. Cited answers, role boundaries, privacy checks and audit correlation.Read the transcript

Delivery status

Shipped is not the same as planned.

Running in production

Platform gateway, two RAG assistants, SSO, deployment isolation, operational logging, audit transport and evaluation workflow.

In progress

  • Dify-based agent composition and platform-level guardrail policy
  • SIEM detection rules for injection attempts and policy violations
  • TLS for the dedicated audit channel
  • Warmup and concurrency control for on-prem inference under load
Labs & experiments

Selected labs.

04 public builds · 2026
Click a lab to open the story
cover — Neural Bonsai01 — AI / MLLiveView case ↗
Neural Bonsai2026 · Live

A browser-only AI playground where model behavior becomes inspectable: draw two-class data, train a real TypeScript MLP, prune weights, and ask a bilingual tutor what changed.

cover — Generative Glyphs02 — FrontendLiveView case ↗
Generative Glyphs2026 · Live

A live typography lab where a generative system becomes a usable product surface: grow glyphs, branch like git, switch renderers, and export shareable artifacts.

cover — ApsisDB03 — BackendLibraryView case ↗
ApsisDB2026 · Library

An embedded Go database for cyclic telemetry: orbit-aware storage, phase queries, deviation checks, rollups, and SQLite-backed durability for edge or local workloads.

cover — Aurora Shader Lab04 — ExperimentLabView case ↗
Aurora Shader Lab2026 · Lab

A fullscreen GLSL workbench for procedural auroras, caustics, and crystal lattices, with live in-browser editing and instant shader reloads.

Contact

Ready for the
next system.

Open to full-time AI platform, LLM infrastructure and LLMOps roles — especially where data sensitivity, reliability and auditability matter.

hello@mareevapps.com
Local time—:—
AvailabilityOpen to full-time roles · remote or relocation
CurrentlyGuardrail policy · audit rules · inference under load
LanguagesRussian · English B2
GitHub ↗Download CV ↓
© 2026 George (Gosha) Mareev(=^.^=)
Open