Skip to main navigation Skip to search Skip to main content

Responsible generation and evaluation of agentic AI: RAIDT as run-level evidence for human oversight and contestability

Research output: Chapter in Book/Report/Conference proceedingConference contribution

3 Downloads (Pure)

Abstract

While agentic AI extends generative AI from producing a single response to completing a multi-step trajectory - retrieve information, call tools, revise intermediate outputs and initiate actions, it also creates a human-computer interaction problem as well as a governance challenge: an evaluator who was not present during execution must be able to reconstruct what the agent did, what evidence it used, which controls were active and how human judgement entered the process. This design-theory guided study identifies the run-level evidence needed for retrospective oversight, organises that evidence as a human-facing review interface, and assesses evidence readiness separately from domain correctness and safety. The paper extends RAIDT - Responsibility, Auditability, Interpretability, Dependability and Traceability - to configured agentic runs through a toolkit development. The result is a minimum evidence pack, anchored 1-5 scoring criteria and a reviewer interaction sequence, which is demonstrated in three scenarios showing transferability across different risk contexts. The contribution is a transparent design account of run-level evidence as an interface for reconstruction, contestation and calibrated reliance in hybrid human-agent systems.
Original languageEnglish
Title of host publicationThe 4th International Conference on Artificial Intelligence and Human-Computer Interaction (ArtInHCI2026), Wuhan, China, from October 16-18, 2026
PublisherIOS Press
Number of pages12
Publication statusAccepted for publication - 7 Jul 2026

Publication series

NameFrontiers in Artificial Intelligence and Applications (FAIA)
ISSN (Print)0922-6389
ISSN (Electronic)1879-8314

Keywords

  • Agentic AI
  • Human oversight
  • Human-AI interaction
  • AI governance
  • Run-level evidence
  • RAIDT
  • Contestability
  • Auditability
  • Traceability

Fingerprint

Dive into the research topics of 'Responsible generation and evaluation of agentic AI: RAIDT as run-level evidence for human oversight and contestability'. Together they form a unique fingerprint.

Cite this