
Tomás Ferreyra
Senior LLM Engineer
Córdoba, Argentina · UTC-3 (America/Argentina/Cordoba) · 7 years of experience
Tomás focuses on measurable model behavior, reliable retrieval, and production evaluation.
Request a shortlistWhy this profile fits
Evaluation before optimization
Defines datasets, graders, and regression checks so model changes are measured against real failure cases.
Reliable retrieval and tool use
Improves grounding, structured outputs, tool permissions, and recovery behavior for production workflows.
Cost and latency control
Treats task success, response time, and model spend as one production system rather than separate concerns.
Relevant toolkit
Delivery evidence
Selected work
Per-Request Model Router & Cost Governor
B2B support-automation SaaSDesigned a routing layer that scores each request and picks the cheapest model (Claude, GPT, or Gemini) that clears a Braintrust quality gate, with semantic caching and prompt compression on top. Cut monthly inference spend 64 percent and p95 latency by more than half while holding CSAT flat.
Eval-Gated Prompt CI
Fintech lending platformBuilt a CI pipeline where every prompt or model change runs an offline eval suite plus a 5 percent canary before promotion, with results posted to pull requests. Regressions caught pre-production rose to 9 in 10, and the team shipped prompt changes daily instead of weekly.
Relevant experience
Career timeline
Senior LLM Engineer
Jan 2023–PresentCorven (fintech payments scale-up) · Remote (US client)
- Designed a per-request model router across Claude, GPT, and Gemini that moved 71 percent of traffic to smaller models and cut monthly inference spend from 48k to 17k US dollars with no measurable quality loss.
Machine Learning Engineer (Applied NLP)
Feb 2021–Jan 2023Helio (B2B customer-support SaaS) · Remote
- Migrated an intent-classification stack from fine-tuned BERT to few-shot LLM prompting, cutting model-maintenance time 60 percent while improving F1 from 0.83 to 0.90.
Machine Learning / Data Engineer
Jun 2019–Feb 2021Datalexa (e-commerce analytics vendor) · Córdoba, Argentina (hybrid)
- Built recommendation and churn models in Python and scikit-learn that lifted repeat-purchase rate 11 percent for a mid-market retail client.
Hiring this role?
See how Tomás fits the LLM Engineer role.
Skill depth
Show experience in context.
Relevant experience by skill
See the complete skill index
Languages
Frameworks
Libraries/APIs
Tools
Paradigms
Platforms
Storage
Other
Education
Formal background
Licenciatura en Ciencias de la Computación, Computer Science
Universidad Nacional de Córdoba (UNC) · 2013–2018 · Graduated with honors; thesis on neural sequence models for text
Especialización en Inteligencia Artificial, Artificial Intelligence
Instituto Tecnológico de Buenos Aires (ITBA) · 2020–2021
Credentials
Certifications
Google Cloud Professional Machine Learning Engineer
Google Cloud · 2024 · Certified
NVIDIA-Certified Professional: Generative AI LLMs
NVIDIA · 2026 · Certified
Communication
Languages
Spanish
Native
English
Fluent
Portuguese
Conversational
More examples
Similar LLM Engineer profiles.

Aravind Balakrishnan
Senior LLM Engineer
Senior LLM engineer, 13 years, specializing in RAG retrieval quality and eval-driven model behavior.

Kacper Wiśniewski
Senior LLM Engineer
Senior LLM engineer who makes model behavior measurable, not anecdotal.

Henrique Vasconcelos
Senior LLM Engineer
Senior LLM engineer, 9 years, specializing in structured outputs and tool calling.