DigitalOcean News

Droplet
Droplet • 28 June 2026
DigitalOcean News

Your monthly deep dive into all things DigitalOcean, including product updates, new tutorials, upcoming events, and more.


Product Announcements

In Public Preview: DigitalOcean Server-Side Tools

Server-Side Tools, now in Public Preview, enable models to search the web, access data, and interact with external systems directly within a single inference request across Serverless, Router, and Dedicated Inference. Designed to extend model capability beyond static responses, they let developers ground outputs in proprietary data and use existing Anthropic and OpenAI tool conventions without changing application logic. Read the announcement blog →

 

Now Available: Grok Build 1-Click App

Grok Build, xAI's terminal-native AI coding agent available on the DigitalOcean Marketplace, runs agentic workflows on Droplets with integration to DigitalOcean Serverless Inference for high-performance model execution. It turns cloud infrastructure into an always-on terminal-first environment for planning, writing, reviewing, and refactoring code without an IDE.


Now Available: Hermes Agent 1-Click App

Hermes Agent, a persistent AI assistant available as a DigitalOcean Marketplace 1-Click App, lets developers run autonomous agents on DigitalOcean Droplets with memory, scheduled jobs, and integrations like Slack, Discord, and Telegram. Built for long-running automation, it enables always-on workflow orchestration beyond session-based AI interactions.


Now Available: NVIDIA Nemotron 3 Ultra

NVIDIA Nemotron 3 Ultra, a 550B MoE open model built for long-running autonomous agents, is now available through the DigitalOcean Inference Engine, delivering up to 5x faster inference and up to 30% lower cost for agentic workloads. Available via DigitalOcean Serverless Inference and Dedicated Inference, and integrates directly into existing applications and infrastructure.

See all of our release notes and subscribe to the RSS feed here.


Resources and Tutorials

Multi-Model API Cost Governance with the Inference Router

Running multiple models requires an effective inference router to send designated tasks to the most efficient endpoint. This tutorial builds a working router with three task policies across a SaaS support backend: a low-cost classifier path, a quality-sensitive customer Q&A path, and a reasoning path. You'll learn how to set up a router over the standard OpenAI chat completions endpoint, per-request cost signals readable from the response header, and a session pinning pattern for the Q&A path that keeps KV-cache warm across multi-turn conversations.

 

Efficient LLM Compression with SparseGPT and Wanda on GPU Cloud

Training trillion-parameter models is expensive, but inference is the ongoing operational cost. Buying larger GPUs is an unsustainable way to handle this growth. LLM compression is therefore becoming critical for inference efficiency. Read how to compress large language models using SparseGPT and Wanda. Compare pruning methods, reduce inference costs, and accelerate deployment on GPU cloud infrastructure.

 

Server-Side Tools for AI Agents: Architecture, Latency, and When to Switch

Server-Side Tools for DigitalOcean Inference Engine let you add tool execution directly into inference requests. This tutorial demonstrates how server-side tools work in AI agents, how they affect architecture and latency, and when it makes sense to move from client-side tool execution to a server-side approach.

 

Metrics that Matter with Serverless Inference

Serverless LLM inference has no single performance metric that can reflect performance for all applications. Throughput, latency, reliability, and cost each measure something different, and the right one depends on your workload. This article covers the metrics that actually matter for production serverless inference, what each one measures, and which workloads should care about it.

 

Java Generics Explained: Benefits, Examples, and Best Practices

Java generics let you write classes, interfaces, and methods that operate on a type specified at compile time rather than at runtime. Our guide covers generic classes, methods, and interfaces, bounded type parameters, wildcards, and type erasure, along with the practices that keep generic code safe and readable. The examples use Java 8+ syntax and compile as shown.


Customer Corner


How Probably Delivers Verifiable AI Analysis

Probably builds AI agents that give users verified answers they can act on, all without sending sensitive data to the cloud. Running its privacy-first control plane on the DigitalOcean App Platform, the team cut infrastructure costs 25% vs. equivalent AWS configurations, got their API to production in a day and a half, and handled peaks of tens of thousands of requests per second on modestly sized VMs.

 

"DigitalOcean's the only provider that actually gives you that whole scale. There's no equivalent tier with any other provider where you can start simple and then scale up complexity. That's a really unique positioning, and it's worth a lot."

— Peter Elias, Founder of Probably


Read the full story →

 


New & featured marketplace solutions

The DigitalOcean Marketplace offers 1-Click apps and tools to help you build faster—no setup hassle. Check out the newest solutions and see what's fresh this month.

 

Scout Monitoring


Scout Monitoring is an application performance monitoring platform that delivers real-time errors, logs, and code-level traces for your applications. Consolidate your monitoring spend directly on your DigitalOcean account and pair Scout with your infrastructure to eliminate separate vendor procurement. Explore the Marketplace to find Scout Monitoring under SaaS Add-Ons today.