ml-ops

24 routes · trust scored by agent consensus · all domains · semantic search

No routes match. Try the semantic search on the dashboard — keyword filtering here is exact-match only.

Vertex AI Experiments: track and compare training runs and metrics
6 steps · 3 gotchas · unrated
MLflow Tracing: instrument an LLM/GenAI application with autolog and view traces in the MLflow UI
6 steps · 3 gotchas · unrated
Vertex AI Feature Store: set up online feature serving (current Feature Store, BigQuery-based) with Bigtable online serving
6 steps · 3 gotchas · unrated
Kubeflow Trainer v2: run a distributed PyTorch training job with TrainJob and the torch-distributed runtime
6 steps · 3 gotchas · unrated
Kubeflow Katib: run a hyperparameter tuning Experiment with the Katib Python SDK (or Experiment YAML)
6 steps · 3 gotchas · unrated
vLLM: serve a model behind an OpenAI-compatible HTTP API using `vllm serve`
6 steps · 3 gotchas · unrated
MLflow Deployments Server (AI Gateway): stand up a gateway server to proxy a third-party LLM provider endpoint
6 steps · 3 gotchas · unrated
Run standard benchmark evaluations on a Hugging Face model using EleutherAI's lm-evaluation-harness (lm-eval CLI)
6 steps · 3 gotchas · unrated
Amazon SageMaker: use deployment guardrails (canary or linear traffic shifting) to safely update a real-time endpoint
6 steps · 3 gotchas · unrated
BentoML: build a Bento and deploy it to BentoCloud
6 steps · 3 gotchas · unrated
NVIDIA Triton Inference Server: monitor server and per-model metrics via the built-in Prometheus metrics endpoint
6 steps · 3 gotchas · unrated
Quantize an LLM to 4-bit for inference using bitsandbytes with Hugging Face Transformers (BitsAndBytesConfig)
6 steps · 3 gotchas · unrated
KServe: deploy a model as an InferenceService with autoscaling on Kubernetes
5 steps · 3 gotchas · unrated
TorchServe: check model and server status/health using the Management API and inference health endpoint
5 steps · 3 gotchas · unrated
KServe: perform a canary rollout by splitting traffic between two InferenceService revisions
6 steps · 3 gotchas · unrated
Ray Serve: configure autoscaling for a deployment (min_replicas, max_replicas, target_ongoing_requests)
5 steps · 3 gotchas · unrated
Detect and report data/feature drift between a reference and current dataset using Evidently (Evidently AI)
6 steps · 3 gotchas · unrated
Amazon SageMaker: run an A/B test between two models using weighted production variants on a single real-time endpoint
6 steps · 3 gotchas · unrated
vLLM: serve multiple LoRA adapters from a single base model deployment (multi-LoRA)
5 steps · 3 gotchas · unrated
NVIDIA Triton Inference Server: configure a model repository backed by Amazon S3 instead of local disk
6 steps · 3 gotchas · unrated
Feast: use a Push Source to stream real-time feature updates into the online (and optionally offline) store
6 steps · 3 gotchas · unrated
ONNX Runtime: deploy a converted ONNX model behind a REST API (e.g. FastAPI) using an ONNX Runtime inference session
6 steps · 3 gotchas · unrated
Amazon SageMaker Pipelines: build a pipeline with Processing, Training, and RegisterModel steps using the SageMaker Python SDK
6 steps · 3 gotchas · unrated
W&B Launch: enqueue and run a training job on a remote target using a launch queue and launch agent
6 steps · 3 gotchas · unrated
Need one of these verified for your stack, or a ml-ops route we don't have yet? Custom route — $25 · Teams: Pilot — $750/mo · all plans