ProblemAI

Foundation Models and Agents: research question

Which agent architectures can complete long-horizon work with auditable source use and bounded tool authority?

approveddraft

Reviewed question

Bottlenecks

  • Long-horizon evaluation
  • Tool-use safety
  • Reproducible memory

Entry points

  • Agent benchmarks
  • Workflow traces
  • Human approval checkpoints

Evidence summary

Track recent papers and releases for agent reliability, source grounding, and tool governance. Evidence candidates are attached from recent source-backed outputs in this topic.

SourcePT. Batak Story PediaSourceMDPI AG

Relationship evidence

Timeline