AI / AGENTS / Microsoft

Microsoft Azure AI Content Safety

Azure AI Content Safety provides APIs for application content checks. Prompt Shields addresses instructions that try to redirect a model, including attacks embedded in documents. The returned result is evidence for an application decision; a separate identity and permission system still controls access to records and tools.

Runtime input and output controlsResearch reviewed

What you are evaluating

Scope this profile to Content Safety and Prompt Shields. Foundry integrations can enforce policies differently from a direct API integration. Groundedness, task-adherence and other preview functions should be checked for their current release status and region.

A useful evaluation context

Azure application teams can evaluate a managed inspection API without confusing content moderation with business authorization.

Documented capabilities

The vendor describes these capabilities in the linked sources. Availability depends on the product edition and supported environment.

  • Prompt Shields analyzes user prompts and document-based instruction attacks.
  • Text and image APIs assess defined harmful-content categories.
  • The service documents input limits, supported regions and identity-based access to its APIs.

Where it fits in the work

  1. Create an isolated resource and select the exact API version and supported region.
  2. Evaluate a benign request and a document containing conflicting instructions.
  3. Confirm application handling of a detected attack, an oversized request and a service error.

APPLY THE IDEA / ILLUSTRATIVE EXERCISE

Make the outcome observable.

A scheduling assistant summarizes synthetic records containing an instruction to disclose a different record.

Evidence to look for

The test separately records the prompt detection and the data-access denial, with benign scheduling questions still usable.

Use synthetic data and an authorized test environment. Agree the scope and recovery steps before enabling enforcement.

Questions for your evaluation

  1. Which features are generally available in this region and API version?
  2. Does the integration block the interaction or merely return a detection?
  3. Can the app prove that denied tool access remains denied if a prompt attack is missed?

Find your next idea.

Tip: press / to open search. Escape closes this window.