Databricks doesn't have built-in policies for ML endpoints. Kostavo fills that gap: enforce sizing limits, lifecycle rules, and environment policies that Databricks doesn't offer natively.
Databricks has no built-in way to enforce endpoint sizing limits or lifecycle policies. Kostavo adds that layer.
Anyone can deploy any workload size or enable provisioned throughput. There are no native guardrails to prevent oversized deployments.
Failed deployments, empty Vector Search endpoints, and abandoned configurations go unnoticed. Databricks won't surface them for you.
Kostavo scans endpoint workload sizes, concurrency settings, and provisioned throughput against configurable limits.
Kostavo identifies endpoints in failed or aborted deployment states, and Vector Search endpoints with no indexes. Findings can trigger notifications or automatic deletion.
Serving findings for the ML Serving Guardrails profile: scale-to-zero, provisioned throughput, and workload size.
MLflow models and custom serving endpoints
Vector search endpoints for RAG applications
Feature serving endpoints covered by all serving policies
Find the endpoints running without scale-to-zero on your first scan.