Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 30 tools.
No remote MCP endpoint is published for this listing.
No capability manifest has been published for this listing yet.