Skip to main content
Back to case studies
Erik GafniOCT 03, 20244 min read

How Sanas achieved a 40% faster ML dev cycle with MLOps and ETL

Erik GafniOCT 03, 20244 min read
40%
Faster dev cycle
70%
Lower infra cost

Share:

How Sanas achieved 40% faster ML Dev cycle with MLOps and ETL

The challenge

As Sanas gained traction and more clients, the AI team needed to scale on every axis at once: enhance the technology, improve development processes, and optimize the entire machine learning workflow to handle a higher volume of requests and data. To get there, the team brought in three experienced Eventum consultants full-time.

Team expansion and best practices

The first step was growing the team beyond its initial seven members. Eventum worked with the existing team to assess strengths and gaps, then recruited ML engineers and data scientists against the specific skills that were missing. To keep a growing team consistent, we introduced a set of shared engineering practices: the team adopted pytorch_lightning, hydra, unit testing, wandb, poetry, linters, formatters, and typecheckers to streamline development. The result was a more efficient and reliable workflow, with fewer errors and less time lost to debugging.

Infrastructure and CI/CD improvements

To meet increasing computational demands, the team rebuilt its CI/CD around the realities of its production environment: pipelines ran on GPUs and Windows machines that mirrored production, so changes were validated quickly and issues surfaced early in the development cycle. Eventum also worked with Sanas to adopt an autoscaling Kubernetes cluster for the dynamic workload of ML tasks, scaling resources automatically with demand and keeping utilization efficient through peak periods.

Cost optimization and cloud migration

Sanas implemented cost-saving strategies to make its AI operations economically sustainable. Enabling spot instances for non-time-sensitive training tasks cut those costs by up to 70% by using otherwise idle cloud capacity. The team also migrated operations to AWS, gaining elastic infrastructure and room to grow into future demand.

Deployment and ETL

Deployment moved to automatic releases via git tags, reducing manual errors and making rollouts routine. Dockerized repositories packaged applications and dependencies consistently, so every stage of the lifecycle ran in the same environment. For ETL, we recommended Dagster to manage data workflows, and Sanas became one of Dagster's first official Trusted ML Partners.

Merge-request development environments

One of the most consequential changes was automatic merge-request development environments: every MR gets an environment where any ETL step can run from any point in the graph. Developers stopped waiting on full pipeline runs to validate their piece of the work, which sped up development dramatically.

The result

A 40% faster ML development cycle. Up to 70% lower training costs on spot-eligible workloads. And a team that had been stuck at seven with the practices, infrastructure, and data pipeline to keep growing without slowing down.

Conclusion

Scaling an AI team is more than hiring. With three Eventum consultants embedded full-time, Sanas expanded the team, modernized its MLOps, optimized its infrastructure and cloud spend, and rebuilt its ETL on Dagster, all in support of its accent-translation and voice-enhancement products. The merge-request environments alone changed how fast the team could move; the whole program positioned Sanas to keep shipping as it grows.

Summarize with AI:

ChatGPTGrokGeminiClaude
Related service

Discovery, architecture, build, evaluation, deployment, handoff. Senior technical ownership end-to-end.

Diagonal halftone representing the AI project delivery flow.