Erik Lund, ai infrastructure engineer

Erik Lund

Senior AI Infrastructure Engineer

Stockholm, Sweden · Europe/Stockholm (UTC+2) · 13 years of experience

Erik focuses on reliable, observable, and cost-efficient compute platforms for AI workloads.

Request a shortlist

Why this profile fits

Reliable AI compute

Runs Kubernetes and GPU platforms with workload isolation, SLOs, capacity controls, and recovery procedures.

Infrastructure as code

Makes environments repeatable through reviewed modules, GitOps, policy checks, and controlled change.

Cost and capacity discipline

Connects utilization, latency, availability, and unit economics to practical scaling decisions.

Relevant toolkit

KubernetesTerraformNVIDIA GPUsPrometheusArgo CD

Delivery evidence

Selected work

Distributed Training Cluster

autonomous systems

Designed and delivered a AI compute platform focused on high-performance training clusters; reduced distributed training interruptions by 64% through topology-aware scheduling and recovery. The release included capacity controls, workload isolation, SLOs, disaster recovery, and unit-economics reporting.

KubernetesTerraform

GPU Failure Recovery Controller

autonomous systems operations

Built the supporting control and measurement layer for the primary system, covering review queues, regression checks, operational visibility, and a documented handoff to the owning team.

KubernetesLinux
View more profile detail →

Relevant experience

Career timeline

Senior AI Infrastructure Engineer

2022–Present

Devlyn client assignments · Remote

  • Led high-performance training clusters delivery for a autonomous systems team and reduced distributed training interruptions by 64% through topology-aware scheduling and recovery.

AI Infrastructure Engineer

2012–2021

autonomous systems product company (confidential) · Stockholm, Sweden

  • Built production systems in autonomous systems, with increasing ownership of reliability, testing, and stakeholder delivery.

Skill depth

Show experience in context.

VerbalCommunicationDomainUnderstandingProblemSolvingCodeQualitySystemDesignDeliverySpeed
Competency shapeAssessment across the same six dimensions used for every profile.

Relevant experience by skill

Kubernetes13 years
Terraform12 years
NVIDIA GPUs11 years
Linux10 years
Prometheus13 years
Argo CD12 years
See the complete skill index

Languages

PythonTypeScriptSQLBash

Frameworks

TerraformNVIDIA GPUsLinuxHelmAnsiblePython

Tools

OpenTelemetryKarpenterVaultGitHub ActionsDocker

Platforms

AWS EKSGoogle Kubernetes EngineAzure Kubernetes ServiceNVIDIA DGX

Storage

S3CephNVMePostgreSQL

Paradigms

GPU schedulingInfrastructure as codeSRECapacity engineering

Credentials

Certifications

NVIDIA-Certified Professional: AI Operations

NVIDIA · 2026 · Certified

AWS Certified DevOps Engineer – Professional

Amazon Web Services · 2026 · Certified

Communication

Languages

English

Fluent

More examples

Similar AI Infrastructure Engineer profiles.