Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
Add this tool to a workspace, then choose which people and agents can use it.
Open workspaceNo capability manifest has been published for this listing yet.