Kacper Wiśniewski, llm engineer

Kacper Wiśniewski

Senior LLM Engineer

Wrocław, Poland · UTC+1 · 11 years of experience

Kacper focuses on measurable model behavior, reliable retrieval, and production evaluation.

Request a shortlist

Why this profile fits

Evaluation before optimization

Defines datasets, graders, and regression checks so model changes are measured against real failure cases.

Reliable retrieval and tool use

Improves grounding, structured outputs, tool permissions, and recovery behavior for production workflows.

Cost and latency control

Treats task success, response time, and model spend as one production system rather than separate concerns.

Relevant toolkit

PythonLangGraphClaude APIOpenAI GPT APIPostgreSQL

Delivery evidence

Selected work

Eval Harness & Regression Gate

B2B SaaS support automation

Designed a golden-dataset eval harness with LLM-as-judge rubrics calibrated to human labels, wired into CI to block any prompt or model change that regressed quality. Cut hallucinated answers 58% and eliminated post-deploy rollbacks over two quarters.

PythonBraintrust

RAG Evaluation Framework

Legal-tech document search

Built a retrieval evaluation framework with labeled query sets and automated recall, precision, and faithfulness scoring on every index change. Surfaced chunking and embedding regressions early and lifted answer relevance (NDCG@10) by 27%.

Pythonpgvector
View more profile detail →

Relevant experience

Career timeline

Lead LLM Evaluation Engineer

2023–Present

Confidential AI Developer Platform (Series B) · Wrocław, Poland (Remote)

  • Designed and owned the company-wide LLM evaluation platform on Braintrust and LangSmith, covering 30+ features and gating every prompt, model, and RAG change behind automated CI quality checks.

Senior Machine Learning Engineer

2020–2023

B2B SaaS Search Scale-up · Wrocław, Poland

  • Led migration of keyword search to a RAG system over pgvector, lifting answer relevance (NDCG@10) 27% and deflecting 45% of tier-1 support conversations.

Machine Learning / NLP Engineer

2017–2020

Fintech Analytics Scale-up · Kraków, Poland

  • Built NLP classification and entity-extraction pipelines for financial documents, reaching 94% F1 and automating a process that previously took analysts 12 hours per batch.

Skill depth

Show experience in context.

VerbalCommunicationDomainUnderstandingProblemSolvingCodeQualitySystemDesignDeliverySpeed
Competency shapeAssessment across the same six dimensions used for every profile.

Relevant experience by skill

Python11 years
Applied ML & NLP9 years
FastAPI7 years
LLM Evaluation5 years
Prompt Engineering5 years
RAG & Vector Search4 years
LangGraph / LangChain4 years
Tool Calling & Structured Outputs4 years
See the complete skill index

Languages

PythonTypeScriptSQLBashGo

Frameworks

FastAPILangGraphLangChainPydanticDSPyRay Serve

Libraries/APIs

Anthropic Claude APIOpenAI APIGoogle Gemini APIOpenAI EvalsGuardrailsInstructorLiteLLMHugging Face Transformerssentence-transformers

Tools

BraintrustLangSmithWeights & BiasespytestGitHub ActionsDockerGrafanaPromptfoo

Paradigms

Eval-driven developmentRetrieval-augmented generationLLM-as-judgeTool / function callingStructured outputsPrompt engineeringRegression testingA/B experimentation

Platforms

AWSGCPAzure OpenAIModalKubernetes

Storage

pgvectorPineconePostgreSQLRedisQdrantWeaviateS3

Other

OpenTelemetryMLflowCost & latency optimizationPrompt versioningGolden datasetsHuman label calibration

Education

Formal background

MSc, Computer Science (Machine Learning & NLP)

AGH University of Science and Technology, Kraków · 2013–2015 · Graduated with distinction

BSc, Computer Science

Wrocław University of Science and Technology · 2010–2013

Credentials

Certifications

Databricks Certified Generative AI Engineer Associate

Databricks · 2022 · Certified

Google Cloud Professional Machine Learning Engineer

Google Cloud · 2025 · Certified

Communication

Languages

Polish

Native

English

Fluent

German

Conversational

More examples

Similar LLM Engineer profiles.