Articles from Source: DigitalOcean-Blog

Intelligence is yours. Let's keep it that way.

2026-09-14 18:43
Intelligence was once solely owned by founders, but that landscape is shifting. Now, powerful providers can control the intelligence layer, risking concentration of power. Open technologies, like open weight models and harnesses, are driving an Open Intelligence movement. This aims to restore ownership and transparency to founders, allowing them to build freely without dependency on single providers. Join us at the Open Intelligence Summit in San Francisco on October 12-13 to discuss the...
Source: DigitalOcean Blog
Paddy Srinivasan

Built for agents: Omarchy's pipeline moves to DigitalOcean

2026-09-09 15:30
🚀 Omarchy, the innovative keyboard-first Linux desktop by David Heinemeier Hansson, has migrated its production pipeline to DigitalOcean. This move enhances their ability to manage bursty agent workloads efficiently. DigitalOcean’s infrastructure supports a streamlined process for provisioning Droplets, executing tasks, and optimizing resources. Additionally, DigitalOcean joins the Omacom Foundation as a Founding Corporate Patron, reinforcing their commitment to open source. 🌐 Stay tuned for...
Source: DigitalOcean Blog

Introducing v5 Droplets: next-generation performance, sized to your workload

2026-08-26 02:19
🚀 Introducing v5 Droplets! Built on 5th Gen AMD EPYC™ processors, these new compute options offer up to 30% higher performance for demanding workloads like AI platforms and high-traffic web applications. 🔧 Customize your Droplet with independent selections for vCPU, memory, and storage, tailoring it to your needs. 💡 Existing Droplet plans remain unchanged, ensuring familiarity and simplicity for users. 🌍 Available now in select regions, v5 Droplets support ambitious teams in scaling their AI...
Source: DigitalOcean Blog
Krishna Nallamothu

Private Preview: DigitalOcean Managed Agents Runtime Services

2026-08-25 19:01
🚀 DigitalOcean has launched Managed Agents Runtime Services (M.A.R.S.) in Private Preview! This new service allows developers to run coding agents and workflows in a fully managed environment. It combines Harness Runtime for execution and Action Gateway for secure access to tools like GitHub and Jira. M.A.R.S. offers features such as durable sessions, isolated execution in Firecracker microVMs, and centralized governance for agent interactions. Developers can maintain flexibility without...
Source: DigitalOcean Blog
Salman Paracha

Patching at Fleet Scale, Twice: How DigitalOcean Closed Januscape and the AMD Safe RET Issue Without Customer Impact

2026-08-24 21:25
In July, DigitalOcean faced two significant security vulnerabilities: Januscape and AMD Safe RET. The Januscape flaw was patched fleet-wide in just eight days, with zero customer impact. This involved livepatching and careful rollout strategies. Shortly after, a separate AMD issue required a different approach, impacting 1,600 hypervisors. The team efficiently managed the situation through structured coordination and automation, achieving full remediation in less than 1.5 weeks. DigitalOcean...
Source: DigitalOcean Blog
Tim Lisko

DigitalOcean Inference Router, Now Cache-Aware: Why the Cheapest Model Isn't Always the Best Deal

2026-08-20 21:20
🚀 DigitalOcean's Inference Router has become cache-aware, addressing the challenge many companies face in managing AI costs as usage increases. This enhancement optimizes the entire agentic session by considering cached context, rather than just selecting the right model. Caching can significantly reduce costs and latency, making it a crucial strategy for developers. 💡 Explore how improved routing can enhance your AI workflows! #DigitalOcean #AI #InferenceRouter #Caching #TechInnovation
Source: DigitalOcean Blog
Salman Paracha

Under the Hood: Serving Kimi K3

2026-07-30 17:10
🚀 DigitalOcean has launched Kimi K3, quickly becoming a top model on the platform! With impressive stats—2.78 trillion parameters and optimized hardware from NVIDIA and AMD—K3 is designed for high performance. The team collaborated extensively to ensure smooth integration and robust verification against benchmarks, leading to enhanced user experience. Explore K3 today on DigitalOcean's Inference Engine! 🌐✨ #KimiK3 #DigitalOcean #AIInnovation #GPU #TechNews
Source: DigitalOcean Blog
Shree Murthy

Outperforming Fable 5 at half the price: meet model synthesis, a new server-side tool on DigitalOcean Inference Engine

2026-07-23 20:03
🚀 DigitalOcean has launched Model Synthesis, a new tool in its Inference Engine. This tool enhances AI performance by combining outputs from multiple models. In tests, the GLM 5.2 + Kimi K2.6 panel scored 65.65% quality at just $0.83 per task, outperforming Fable 5, which scored 62.21% at $1.59. Model Synthesis is now available in Public Preview, allowing users to choose optimized presets or define their own configurations. #DigitalOcean #AI #ModelSynthesis #InferenceEngine #TechNews
Source: DigitalOcean Blog
Tyler Gillam

Upcoming GPU Pricing Updates

2026-07-21 00:30
🚨 Price Update Alert! 🚨 Starting August 1, 2026, DigitalOcean will adjust prices for select on-demand NVIDIA and AMD GPU droplets. Any active workloads will be billed at the new rates effective September 1, 2026. For those on 12-month reserved plans, current rates remain unchanged until contract renewal. For details, visit our pricing page! #DigitalOcean #GPUPricing #CloudComputing #TechUpdates
Source: DigitalOcean Blog

Scale Faster with Managed Weaviate: Now in Public Preview on DigitalOcean

2026-07-09 19:08
🚀 Exciting news for developers! Managed Weaviate is now in public preview on DigitalOcean. This service allows you to run Weaviate in production effortlessly, starting at just $20/month. It handles backups, security, and scaling, letting you focus on building your applications. Enjoy predictable pricing with no hidden fees, and maintain full compatibility with existing Weaviate clients. Get started today and streamline your AI solutions! 🌐🔍 #Weaviate #DigitalOcean #AIDevelopment...
Source: DigitalOcean Blog
Waverly Swinton

Built for Mass Scale: Hard-Won Lessons from Teams Running High Volume Inference Workloads in Production

2026-07-02 10:00
🚀 Transitioning AI from prototypes to high-volume production comes with significant challenges. At DigitalOcean Deploy 2026, leaders from Workato, Hippocratic AI, and ISMG shared crucial lessons on managing latency, security, and infrastructure. 🔍 Key insights included the importance of policy-aware systems and agent permissions to enhance reliability. 💡 The takeaway? Scaling AI is about smart architecture and management, not just model performance. #AIEvolution #DigitalOcean #TechTalk...
Source: DigitalOcean Blog
Hasan Nabulsi

DigitalOcean Evaluations: Production Model and Router Testing for the Inference Stack

2026-07-01 15:41
🚀 DigitalOcean introduces Evaluations on its Inference Engine, allowing teams to validate model and router configurations with their own data before production. Key features include: - LLM-as-a-Judge scoring for performance tracking. - Six pre-built metrics and custom rubrics for tailored evaluations. - Evaluation presets to save configurations for easy reuse. Efficient dataset management and programmatic triggers streamline the evaluation process. Start validating your models today! 🌐...
Source: DigitalOcean Blog
Grace Morgan

Run Codex in the cloud – DigitalOcean for Codex is now available

2026-06-25 21:16
🚀 Exciting news for developers! The DigitalOcean plugin for Codex is now in Public Preview. This new feature allows you to create and connect Codex-ready cloud development machines directly within Codex, eliminating the need for manual server setup. You can provision a remote development environment using simple commands. 💻🌐 Once connected, you can configure projects, install dependencies, and manage machines—all from your own DigitalOcean account. Plus, with Codex in the ChatGPT mobile app,...
Source: DigitalOcean Blog
Ari Sigal

The Inference Alpha: Maximizing Frontier Models on AMD

2026-06-10 14:27
🚀 At DigitalOcean, we focus on high-performance infrastructure for AI, particularly frontier Large Language Models (LLMs) on AMD GPUs. Our approach emphasizes that peak inference speed is influenced by model architecture and runtime execution, alongside hardware. This "performance alpha" highlights the benefits of specialized inference engineering. Recent collaborations with Wafer demonstrated significant throughput improvements: Kimi 2.5 saw an 11.33x speedup, while DeepSeek V3.2 achieved a...
Source: DigitalOcean Blog
Emilio Andere

What We Learned Hiring 33 Engineers in Two Weeks

2026-06-09 22:58
🚀 Earlier this year, we needed to hire engineers quickly for a product launch. We revamped our interview process to focus on real-world skills instead of outdated methods. 💻 Candidates participated in a hands-on, three-hour build session, utilizing AI tools to prototype solutions. This approach allowed us to evaluate their decision-making and collaboration skills. 🤝 After coding, we engaged in discussions about design choices and real-world challenges, providing a platform for candidates to...
Source: DigitalOcean Blog
Janet Harrah

Model Evaluations: Prove Your Routing Policy Actually Works

2026-06-04 19:52
🚀 Teams often struggle not due to a lack of good models, but because their routing policies falter under real conditions. DigitalOcean's Model Evaluations, now in Public Preview, can help assess models and routing strategies effectively. This tool enables evaluations across cost, latency, and output quality. In the guide, you’ll find steps to compare a single frontier model, an Inference Router, and a Bring Your Own Model (BYOM) on a legal assistant use case. Learn how to set up, run, and...
Source: DigitalOcean Blog
Sathish Jothikumar

The Team Behind Deploy: Shipping AI, the DigitalOcean Way

2026-06-03 19:38
🚀 Deploy 2026 brought together developers, startups, and partners in San Francisco to discuss building and scaling AI products. DigitalOcean unveiled the AI-Native Cloud, featuring over 15 product launches, including the Inference Router. Key sponsors included NVIDIA and MongoDB, while companies like Hippocratic AI showcased their journey. The event highlighted DigitalOcean's culture of ownership, as team members shared insights on customer collaboration and product development. Explore...
Source: DigitalOcean Blog
Sujatha R

Powering the Inference Era: Inside the DigitalOcean Data & Learning Layer

2026-06-03 19:23
Unlock the potential of AI-native applications with DigitalOcean's new Data & Learning Layer! 🌐 This platform integrates structured, vector, and retrieval layers, streamlining development for real-time multimodal pipelines and enterprise knowledge bases. Key features include: - Managed PostgreSQL & MySQL for structured data. - Knowledge Bases for seamless unstructured data management. - Managed Weaviate for vector search capabilities. These tools work together, reducing latency and costs...
Source: DigitalOcean Blog
Spoorthi Rao Nimmala

Open by Design: How NVIDIA and DigitalOcean Are Building the Stack for the Always-On Agentic Era

2026-06-02 18:29
NVIDIA and DigitalOcean recently discussed the evolution of open-source AI at the "Open by Design" session. They emphasized the need for commitment to open models, like NVIDIA's Nemotron, to ensure ongoing improvements for developers. Evaluation standards for AI applications remain a challenge, impacting developers’ confidence. The session also highlighted the importance of sub-agent workflows and effective token economics in scaling AI systems. For more insights, watch the full session! 🎥✨...
Source: DigitalOcean Blog
Jess Lulka

The Inference Tax: How Prefix-Aware Routing Eliminates the Hidden Cost of LLMs at Scale

2026-06-01 19:30
🌐 Inference demand is rising rapidly, projected to dominate AI compute by 2030. A significant portion of compute costs is avoidable due to redundant work in systems. 🔍 DigitalOcean's prefix-aware routing addresses this inefficiency, significantly reducing unnecessary computations. By optimizing GPU performance and caching, they enhance cost-effectiveness without hardware constraints. 🚀 Upcoming improvements in Serverless Inference will make these benefits accessible to all users, ensuring...
Source: DigitalOcean Blog
Simon Mo, CEO of Inferact

DigitalOcean Serverless Inference: A Deep Dive

2026-06-01 18:44
🚀 **Introducing DigitalOcean Serverless Inference!** This API-first platform simplifies AI model deployment at scale. It supports 30+ foundation models across various modalities through a single API key. Key features include automatic scaling, intelligent routing, and built-in tools for efficient model management. Get started easily and pay only for what you use! #DigitalOcean #AI #Serverless #MachineLearning #TechInnovation
Source: DigitalOcean Blog
smehta

AI Disruptors: How the Next Generation of Business is Being Built

2026-05-29 21:30
At the Deploy 2026 conference, I moderated a panel with AI founders discussing what differentiates successful AI products from demos. Key insights included the importance of measuring agent performance and ensuring reliability. Founders like Angela Hoover from Andi AI and Hovsep Seraydarian from LawVo emphasized that human oversight is essential in high-stakes fields. They also highlighted the challenges of model selection in a rapidly evolving landscape, noting that execution and...
Source: DigitalOcean Blog
Dinesh Murthy

OpenCode Now Supports DigitalOcean Inference Router for Intelligent Model Routing

2026-05-28 21:02
🚀 Exciting news for developers! DigitalOcean's Inference Router is now available in OpenCode, addressing the costly issue of using a single model for all tasks. This dynamic router intelligently directs requests to the most suitable model, optimizing costs and ensuring efficient resource use. To get started, simply connect your DigitalOcean account in OpenCode and select your Inference Routers. Explore the future of cost-effective AI model routing! 🌐💻 #OpenCode #DigitalOcean #AI...
Source: DigitalOcean Blog
Musa Malik

Scalable, Cost-Efficient AI: Introducing Unified Batch Inference on DigitalOcean

2026-05-27 17:43
🚀 Exciting news from Deploy 2026! DigitalOcean has launched Batch Inference on its AI-Native Cloud, designed for efficient high-volume workloads. This feature allows developers to process up to 100k requests asynchronously at reduced costs, streamlining tasks like data transformation and content generation. With a unified API for OpenAI and Anthropic, managing multiple models is simpler than ever. Batch Inference also helps bypass rate limits, ensuring smoother operations. Explore how this...
Source: DigitalOcean Blog
smehta

Request-Based Autoscaling Is Now Generally Available on App Platform

2026-05-22 18:02
🚀 Request-based autoscaling is now live on DigitalOcean App Platform! Apps can automatically scale based on live HTTP traffic signals like requests per second and P95 response latency. This ensures your infrastructure reacts promptly to user demand. Now available for both shared and dedicated CPU instances, it allows all users to benefit from responsive scaling without needing a plan upgrade. 🔍 Use the Insights tab to understand traffic patterns and configure your autoscaling rules...
Source: DigitalOcean Blog
Greeshma Pillai

How We Built DigitalOcean Inference Router

2026-05-20 14:57
🚀 Exciting news from DigitalOcean! They have launched the Inference Router, designed to optimize model selection for AI tasks. Instead of relying on a single model, this router intelligently routes requests to the most suitable model based on cost, latency, or quality. The Inference Router utilizes a 30B Mixture-of-Experts model, achieving impressive accuracy in task detection. With easy setup via a single line of code, developers can enhance their workflows without the burden of manual...
Source: DigitalOcean Blog
Adil Hafeez

Your Model Doesn't Matter. Your Infrastructure Does.

2026-05-13 16:45
Unlocking the potential of AI starts with the right infrastructure. 🌐 DigitalOcean emphasizes that while everyone has access to similar models, success lies in the surrounding infrastructure—routing logic, data pipelines, and scalable solutions without code rewrites. Their recent session showcased how teams can move seamlessly through serverless, dedicated, and routed setups, maximizing efficiency and reducing costs. 💡 Explore the full capabilities of DigitalOcean's AI platform! #AI...
Source: DigitalOcean Blog
Amit Jotwani

Introducing DigitalOcean AI-Native Cloud for Production AI Workloads

2026-04-28 19:14
🚀 DigitalOcean has introduced its AI-Native Cloud, addressing the growing challenges in AI workloads. The shift in AI infrastructure highlights inference as the new focus, with reasoning models and autonomous agents taking center stage. This full-stack solution simplifies development by reducing complexity, allowing developers to concentrate on building rather than integrating. Key features include the Inference Router, dedicated GPU infrastructure, and a wide range of models available for...
Source: DigitalOcean Blog
Paddy Srinivasan

How we built the most performant DeepSeek V3.2, MiniMax-M2.5 and Qwen 3.5 397B on DigitalOcean NVIDIA HGX™ B300 GPU Droplets

2026-04-28 09:00
🚀 We are excited to announce the launch of DeepSeek V3.2, MiniMax-M2.5, and Qwen 3.5 397B on DigitalOcean Serverless Inference. These models achieve leading performance, with DeepSeek V3.2 delivering 230 output tokens per second and a sub-1-second time to first token for 10,000 input tokens. Fast inference is crucial for modern AI applications, ensuring a seamless user experience. Our optimizations enable businesses to lower costs and maintain high performance. Explore the benchmarks and see...
Source: DigitalOcean Blog
Bhaskar Dutt

DigitalOcean Dedicated Inference: A Technical Deep Dive

2026-04-25 02:51
🚀 DigitalOcean has introduced Dedicated Inference, a managed LLM hosting service designed for teams needing reliable, high-performance inference on dedicated GPUs. It simplifies deployment by handling the orchestration, while users maintain control over model selection and scaling options. This service targets organizations with consistent inference demands, offering predictable costs and performance. Key features include public and private endpoints, Kubernetes-native orchestration, and...
Source: DigitalOcean Blog
dgupta

Beyond the Abyss Project Poseidon’s Quest for Zero-Downtime Reliability

2026-04-23 19:29
🌐 DigitalOcean is advancing its cloud infrastructure with Project Poseidon, aiming for zero-downtime reliability. This new system uses Machine Learning and Generative AI to identify "at-risk" nodes before server crashes occur. By shifting from reactive monitoring to proactive measures, it enhances operational efficiency. The tiered approach filters out 98% of irrelevant data, focusing only on critical signals. Poseidon is designed to evolve continually, ensuring it adapts to new hardware and...
Source: DigitalOcean Blog
Sartaj

From Incident Counting to SLIs: How DigitalOcean Rethought Availability

2026-04-23 09:15
📊 DigitalOcean has redefined its approach to measuring availability by shifting from an incident-based metric to Service Level Indicators (SLIs). Initially, availability numbers fluctuated between 99.5% and 99.9%, often not reflecting true customer experience. The new metric, consistently above 99.95%, better represents actual platform performance. Key changes include separating measurements into Control Plane and Data Plane, allowing for more accurate assessments of service health. This...
Source: DigitalOcean Blog
Miguel Carrera

The LLM Inference Trilemma: Throughput, Latency, Cost

2026-04-22 15:56
Navigating the complexities of Large Language Model (LLM) inference involves understanding the "trilemma" of throughput, latency, and cost. Scaling LLMs isn't as simple as adding more servers; it requires careful optimization. Key cost factors include hardware expenses, electricity, and specialized labor. Each decision impacts the balance between performance and expenses. ⚖️ This comprehensive guide offers insights on optimizing for either throughput or latency, depending on your use case....
Source: DigitalOcean Blog
Balaji Varadarajan

Mastering the 600B+ Frontier: Optimizing Large Model Deployments on the Inference Cloud

2026-04-21 20:10
The landscape of model deployment is evolving rapidly, with weights now exceeding 700GB and parameters reaching trillions. 🧠 Optimizing storage architecture is crucial to combat "Data Gravity," which can slow down GPU performance and increase operational costs. High-bandwidth storage solutions can significantly reduce deployment latency, impacting overall efficiency. 📈 Cloud providers that offer specialized GPU and storage combinations are essential for managing these large models...
Source: DigitalOcean Blog
Brett Snyder

The Inference Cloud Memory Layer: A Technical Dive into DigitalOcean Managed Databases

2026-04-17 20:10
DigitalOcean addresses the growing need for a robust memory layer in AI applications with its Inference Cloud. 🌩️ As AI transitions to production-grade models, the absence of persistent memory can lead to issues like loss of long-term recall and workflow vulnerabilities. DigitalOcean Managed Databases, including PostgreSQL and MongoDB, serve as foundational memory layers to enhance stateful AI applications. This shift to the inference cloud allows developers to focus on building intelligent...
Source: DigitalOcean Blog
Joe Keegan

Load Balancing and Scaling LLM Serving

2026-04-15 19:03
Load balancing for Large Language Models (LLMs) differs significantly from traditional services due to prompt caching. Efficient routing strategies are essential to maximize cache effectiveness and minimize latency. The article explores specialized routers that enhance performance while addressing the limitations of standard load balancing methods. Various inference engines like vLLM and TensorRT streamline the process, allowing for efficient handling of diverse workloads. For optimal...
Source: DigitalOcean Blog
Mohammad Ashar Khan

Building a Robust Documentation Agent with DigitalOcean Gradient AI Platform

2026-04-13 16:59
🚀 At DigitalOcean, we've prioritized documentation by creating an AI assistant that helps developers find answers quickly. This tool allows users to ask questions in plain language and receive accurate, actionable responses. Through extensive testing and validation, we improved the assistant's reliability and performance, ensuring it can effectively guide users. Key components include a robust architecture on the Gradient AI Platform and a focus on metrics for continuous improvement. Explore...
Source: DigitalOcean Blog
Anna Lushnikova

Advanced Prompt Caching at Scale

2026-04-07 19:11
🌐 Prompt caching optimizes inference requests by reusing computed KV states, enhancing efficiency and reducing costs. However, as systems scale with multiple replicas, cache hit rates drop, posing challenges. 🔄 Implementing session affinity can improve performance by routing requests to the same replica, preserving cached data. 📊 Effective architectural strategies, including tiered caching and proper prompt structure, can significantly boost efficiency. #PromptCaching #AIInference...
Source: DigitalOcean Blog
Andrew Dugan

The Hidden Cost of Complex AI Platforms: Why Developer Experience Matters

2026-04-03 15:44
Navigating the cloud AI platform landscape can be challenging. 🖥️ Many developers face significant delays due to unclear documentation, fragmented workflows, and complex setups. Tasks that should take minutes can stretch into hours, impacting productivity and innovation. ⏳ Key factors include the real cost of developer experience, Time-to-First-Value (TTFV), and the hidden complexities of scaling. A seamless integration of tools is essential for faster iterations and successful deployments....
Source: DigitalOcean Blog
Shaoni Mukherjee

The Glue Problem in Modern AI Development

2026-04-02 21:30
AI is transforming software development, yet deploying it remains complex. The challenge lies in integration, where various systems must work together seamlessly. Fragmented setups lead to increased developer effort in maintaining glue code, diverting focus from product features. The article discusses the advantages of a vertically integrated cloud model over a neocloud-hyperscaler combo, highlighting reduced complexity and operational costs. By minimizing integration points, developers can...
Source: DigitalOcean Blog
James Skelton