ricotek engineering
Bleeding-Edge Platform Engineering
Ricotek is a platform engineering consultancy specializing in agentic automation, multi-cloud infrastructure, and autonomous operations.
Ricotek Engineering architects autonomous, event-driven infrastructure and agentic CI/CD pipelines for high-velocity outcomes.
We build robust, multi-cloud platforms where agentic automation drives delivery. From mobile-first control planes to advanced Kubernetes orchestration, we engineer environments where no single person is irreplaceable — because the system runs itself.
core offerings
Agentic systems. Multi-cloud infrastructure. Fully traceable delivery. These aren't separate services — they work together.
Autonomous AI & Agentic Engineering
Custom AI Workspaces
Engineering custom agent gateways with conditional multi-model routing — Anthropic, OpenAI, xAI, DeepSeek, or any provider. Agents are onboarded with role-based access controls, strict guardrails, custom skills and markdown runbooks, and self-learning capabilities.
Headless Operations
Designing mobile-first, chat-integrated operational control planes for infrastructure management, deployment routing, and anomaly triage. Every agent action is fully logged, traceable, and auditable — with token tracking, usage caps, and start/stop controls through a web-based enterprise dashboard.
Automated Delivery Workflows
Deploying multi-agent systems that decompose tickets, review implementations, and guard production deployments — with end-to-end traceability from issue to merge. Integrates with Jira, Linear, GitHub Issues, or whatever your team already uses.
Multi-Cloud Infrastructure & Kubernetes (IaC)
Enterprise-Grade Provisioning
Designing and provisioning full-stack cloud environments — compute, networking, IAM, security boundaries — using infrastructure-as-code. Repeatable, auditable, and version-controlled across AWS, Azure, or whatever cloud you run on.
Kubernetes Mastery
Architecting, upgrading, and hardening production-grade EKS and AKS clusters — multi-cluster failover, zero-downtime version jumps, and full security posture with RBAC and network policies. Custom SRE agents provide round-the-clock observability and self-healing.
Migration Strategy
Provisioning a parallel cloud environment via Terraform and auto-wiring every release from the current environment to the new one. CI/CD silently validates every component and config in the background. By cutover day, both environments are fully equalized — flipping the switch is just a DNS change.
Advanced CI/CD & Platform Ops
Multi-Environment Release Orchestration
Designing complex, multi-environment release orchestrations with operator-authorized pipeline execution.
Headless Agentic IDP
Building agent-driven internal developer platforms where a developer asks for infrastructure and the agent provisions it — no portals, no click-ops, just a conversation.
Deep Observability
Consolidating metrics, traces, and logs across your existing observability stack — whether it's Prometheus, Grafana, Datadog, New Relic, or custom tooling — to ensure high availability across global environments.
trophy case
Real problems, real solutions. Engagements that show how Ricotek delivers.
The AI-Augmented Platform
the problem
Engineering bottlenecks in infrastructure implementation and QA handoffs.
the ricotek solution
Built an agentic delivery workflow using a custom Hermes framework with conditional multi-model routing. Created supervised human-in-the-loop approval gates for safe execution against production-adjacent cloud resources.
Database as a Service
the problem
DBAs needed an Azure DevOps deployment pipeline that could select any build package and push to any database on any host, in any environment — with the capacity for 20 simultaneous deployments running concurrently. Only the build pipeline existed; the deploy side was entirely missing.
the ricotek solution
Built a Python-based web application that dynamically loads every DB host and database via configurable filters, then lets the operator select the build artifact, target hosts, and databases — firing up to 20 concurrent deployments from a single interface.
HBO GO Launch
the problem
Legacy infrastructure bottlenecking the Latin American and Caribbean region launch for HBO Go at HBO Latam HQ. Original timeline projected 2+ years to reach 40 countries.
the ricotek solution
Deployed the entire HBO Go stack onto a hybrid cloud — 400 web servers, multiple SQL clusters, and a Docker Swarm cluster on AWS — with full failover capability between data centers. Built every deployment script and rollback procedure using custom Bash, PowerShell, and Ansible. Shipped 2–3 releases per week and launched all 40 countries within a single year.
HBO was so satisfied they asked Ricotek to provision Cinemax Go as well — this time with a re-architected infrastructure stack using shared web services, centralizing all content and enabling deployments in seconds.
The AI-Augmented Migration
the problem
After two unsuccessful Sisense support calls — each over two hours — the Sisense + AKS upgrade remained stuck. Terabyte-scale volumes needed to be migrated, and each component required 4- to 5-version jumps, demanding multiple sequential upgrades. Doing it manually was a multi-day affair.
the ricotek solution
Fed a phased-out internal markdown runbook to Claude Code and completed the full QA cluster upgrade — Sisense and AKS — in 8 hours. Used Azure cross-region volume replication for the terabyte-scale volumes and Velero for everything else, ending up with both a DR strategy and a migration strategy. For the production cluster, a customized Hermes agent drove the same workflow, finishing in just over 5 hours.
The Autonomous SRE Agent
the problem
Keeping a Sisense Kubernetes cluster healthy meant chasing issues reactively — playing whack-a-mole at all hours. No round-the-clock visibility into cluster health, no automated diagnosis, and no peace of mind.
the ricotek solution
Built an autonomous SRE agent powered by Hermes crons — Python scripts that run 20+ checks across the cluster, collecting metrics and detecting anomalies. When it finds an issue, it publishes to Slack with the discovered problem, root cause, and 2–3 proposed solutions. In many cases it self-heals and simply reports what it fixed. Round-the-clock watch while everyone sleeps.
The Hands-Off Release Pipeline
the problem
A team of 10 woke up at 5 AM for every release. One person was handed an email with 3–10 components, specific versions, and a strict deployment order — then had to manually push buttons, fix issues on the fly, and prep the environment so two QA engineers could run smoke tests on laptops. Everything had to be done by 8 AM.
the ricotek solution
Built a master release project where you pre-select packages, versions, and deploy order — schedule it for 5 AM and walk away. Automated smoke tests run immediately after deploy and results land in your inbox. By 6 AM, everything is done. Team went from 10 to sometimes 3. No more 5 AM wake-ups, and releases can happen any time of day.
The Templated CI/CD Framework
the problem
A company with zero CI/CD — .NET and Angular components were deployed manually via SFTP or file copy. No infrastructure-as-code, no standardized pipelines, and spinning up a new component across environments was a black hole with no timeline.
the ricotek solution
Built a Terraform pipeline framework that provisions container-based web apps and headless servers in Azure, then auto-generates the full build and deploy pipelines. Fully templated in YAML and HCL — anyone on the team, including developers, can spin up a component with CI/CD across all 12 environments. What used to take forever: ticket in the morning, done in a couple of hours. Deployed 15 .NET APIs and one UI this way.
The CRM Mass Deployment Pipeline
the problem
A major CRM component had no deployment pipeline. Two people were handed spreadsheets listing 200+ customers and six different versions — then manually stopped web servers, backed them up, uploaded packages, extracted files, made config changes, and turned sites back on. Rollbacks were messy and ad-hoc.
the ricotek solution
Built a master release pipeline that takes any approved build and pushes it to up to 20 customers simultaneously — push-button, no server logins required. Added a one-click rollback method. Full logging and traceability. One person can now choose a customer, choose a package, hit a button, and walk away — even from a phone or tablet. Anyone on the team can run it.
frequently asked questions
What does a typical Ricotek engagement look like?
We start with an assessment — you bring the backlog, the workflow tools, and the operational constraints. We map the first engagement, show how it gets delivered through your existing systems, and define measurable outcomes. Engagements can be project-based, retainer-based, or embedded — whatever fits your team.
Do you work with our existing tools and workflows?
That's the point. We don't ask you to adopt a new platform. Our agents operate inside your repos, your CI/CD pipelines, your ticketing system (Jira, Linear, GitHub Issues), and your chat tools (Slack, Teams, Telegram). The work appears where your team already works.
How do you handle security and access controls?
Agents are onboarded with role-based access controls, strict guardrails, and full audit logging. Every action is traceable. For sensitive operations, we use human-in-the-loop approval gates — the agent proposes, a human approves. We can operate entirely within your security boundary.
What's the difference between Ricotek and a traditional DevOps consultancy?
Traditional consultancies send you a team of people. We deploy agentic systems that augment your existing team — agents that build, verify, and self-heal around the clock. You get the output of a platform engineering pod without hiring one.
Can your agents operate in our cloud environment?
Yes. We work across AWS, Azure, GCP, and hybrid environments. Agents are scoped to your infrastructure — they don't leave your boundary. We provision with your IaC tooling, deploy through your pipelines, and log to your observability stack.
What kind of teams do you work with?
High-velocity platform, DevOps, and infrastructure teams that are bottlenecked — too much backlog, not enough people, or both. We also work with teams that have no platform engineering function at all and need one stood up from scratch.
How long does it take to see results?
The first engagement typically delivers within the first sprint. For ongoing work, agents operate continuously — you'll see PRs, deployment traces, and status updates flowing through your existing tools from day one.
deep dives
Technical writing on the patterns, architectures, and decisions behind what we build.
Building a Multi-Model Agent Router for Platform Engineering
How we built a conditional routing system that picks the right model for the right task at the right cost — and why one model can't do everything.
how we engage
No black boxes. No long sales cycles. Three ways to work with Ricotek.
Assessment
Bring your platform engineering backlog and constraints. We map the first engagement, identify the highest-impact workflow to automate, and deliver a concrete plan with measurable outcomes. One to two weeks.
Project
Scoped, time-boxed delivery of a specific outcome — a migration, a pipeline overhaul, an SRE agent, a multi-environment deployment system. Fixed scope, clear deliverables, defined timeline.
Retainer
Ongoing platform engineering capacity. Agents run continuously, handling backlog, incidents, optimizations, and operational improvements. You get the output of a pod without hiring one. Monthly, cancel anytime.
Start with an assessment.
Bring the platform engineering backlog, the workflow tools, and the operational constraints. Ricotek maps the engagement and shows how the work gets delivered.