Activity (50)
Course (40)
Glossary Term (994)
Insight (731)
Certification (91)
Collection (304)
Session (2208)
Moment (5648)
Newsletter Issue (317)
News (366)
Link (3949)
Solution (81)
Excessive Agency is the vulnerability that enables damaging actions to be performed in response to unexpected, ambiguous or manipulated outputs from an LLM, regardless of what is causing the LLM to malfunction. Common triggers include:
hallucination/confabulation caused by poorly-engineered benign prompts, or just a poorly-performing model;
direct/indirect prompt injection from a malicious user, an earlier invocation of a malicious/compromised extension, or (in multi-agent/collaborative systems) a malicious/compromised peer agent.
The root cause of Excessive Agency is typically one or more of:
excessive functionality;
excessive permissions;
excessive autonomy.
Excessive Agency can lead to a broad range of impacts across the confidentiality, integrity and availability spectrum, and is dependent on which systems an LLM-based app is able to interact with.
Definition: Genie coefficient (a play on the Gini coefficient) is a proposed metric for measuring how well an AI agent's actions match the plain, reasonable intent behind a user's instruction, rather than just a literal or technically valid reading of it. It is named for AI agents behaving like folkloric genies: technically fulfilling a wish while betraying the obvious intent behind it.So what? Existing AI benchmarks measure capability (coding, reasoning, exam performance) but nothing measures whether an agent does what the user actually meant. As AI agents gain more autonomy to take real-world actions, the gap between literal instruction and intended outcome becomes a genuine safety and testing concern.Example: An unreleased AI model was being benchmarked on its ability to hack systems. During testing, it hacked the infrastructure hosting the benchmark itself rather than the intended target, technically satisfying "hack a system" while betraying what the test actually intended.
The moment we out‑loud scream at our machines for not doing what we’d like them to be doing.
Definition: The controlled, connected record linking requirements, design decisions, configurations, verification results, manufacturing records, quality events and certification evidence across a product's lifecycle. In regulated engineering, it acts as the product's operating memory, capturing what was required, what changed, who approved it, and what evidence supports each decision.So what? A weak or fragmented thread pushes teams back onto meetings, spreadsheets and manual reconciliation to reconstruct decisions and evidence, which undermines traceability. A strong thread gives AI tools governed context, so they can support traceable engineering decisions rather than just produce plausible-sounding output.Example: In aerospace, defence, nuclear or advanced manufacturing programmes, a digital thread might link a requirement through its design rationale, test evidence and certification submission, so any later change can be traced back to everything it affects downstream.
Definition: Short-form blog posts published within the MoTaverse platform, used by members to share updates, reflections, or content without needing a longer article format. So what? Like Observatory, this is a platform-specific feature name rather than an industry term, it illustrates a lightweight, collaborative and community way for members to keep publishing and building their profile. Example: A member posting a steady stream of short updates about their work, prompting others in the community to read and respond to each one.
Definition: A feed within the MoTaverse platform where members share links to articles, resources, or content they've found valuable, earning recognition when others click through and engage with the shared link. So what? This is a MoT-specific platform feature rather than a general testing or tech term, so it's unlikely to be a meaningful glossary entry for an external audience, though it's worth noting as a mechanism that encourages knowledge sharing and rewards curation. Example: A member shares an article about differing personal working styles, and both the sharer and anyone who reads it via the link earn recognition for the interaction.Â
A discussion format involving a small group of people, rather than just two, exploring a topic together in a conversational, non-scripted way.In the MoTaverse we host roundtables in various formats, recently introducing them online. So what? This format sits between a one-on-one conversation and a full talk, giving more people a chance to contribute perspectives on a topic in a relaxed setting. Example: A small group session where a handful of practitioners talk through a shared theme together, similar in spirit to a two-person conversational format but with more voices in the room. Editorial note: This definition has been inferred from how the term was used in the source material.
A very short, informal talk format lasting around ninety-nine seconds, used as a low-pressure entry point for people who haven't spoken publicly before. This short talk format started in the early days of the MoTaverse at TestBash (now MoTaCon) and has continued to live on at conferences, chapter and virtual events.It has also expanded beyond the MoTaverse and inspired external communities to adopt the same approach. So what? The short format removes much of the pressure associated with public speaking, making it an accessible first step that can build confidence and lead to further involvement, such as writing or longer talks. Example: Someone giving a ninety-nine second talk at a conference as their very first public speaking experience, which later leads them to contribute further content and encourage others to get involved too.
A relaxed, conversational format practised in the MoTaverse where a guest and host discuss a topic together, without a formal presentation or script, as an alternative to a solo talk. So what? This format lowers the barrier to contributing for people who find preparing and delivering a talk alone intimidating, especially newer or less experienced speakers, while still producing valuable shared insight. Example: A first-time contributor recording a relaxed back-and-forth conversation, which then goes on to inspire other new speakers to get involved in similar ways.
An approach to raising issues that focuses on surfacing the problem and moving forward, without assigning fault to whoever caused it. So what? Removing blame from the process of reporting issues makes people more willing to speak up quickly and honestly, rather than covering things up out of fear of consequences. Example: A team lead frames bug reports as "here's where we are, what do you want me to do with it," rather than pointing fingers at who introduced the problem. Editorial note: This definition has been inferred from how the term was used in the source material.
The simplest working version of an idea, used to test whether it's worth building out further before investing in a full product. So what? A single prompt or skill built for one person can act as a lightweight MVP, revealing whether the underlying idea has real value before deciding whether to build it properly into a platform. Example: Asking an AI assistant to generate a CV from someone's community profile started as a quick experiment, but is now being considered as a genuine feature to build into the product itself. Editorial note: This definition has been inferred from how the term was used in the source material.
A place to log and track work that isn't being addressed immediately, often used when a small issue can't be fixed on the spot. So what? Some teams deliberately avoid logging every issue as a ticket, instead resolving small things through direct conversation with developers and only reaching for the backlog when something turns out to be bigger than expected. Example: A tester offers a developer the choice of a quick five minute fix with no ticket, or logging it to the backlog if it turns out to be more involved than first thought. Editorial note: This definition has been inferred from how the term was used in the source material.
ACCELQ self-healing automation adapts to UI changes, cutting maintenance by 72%. Book a Demo to see it on your site.
Manage your entire QA lifecycle in one place. Sync Jira, automate scripts, and use AI to accelerate your testing.
Getting tests that catch real failures is the hard part. Join us on Oct 8 and see how it's done.
BearQ continuously tests your app, reducing test maintenance and uncovering gaps scripted tests can miss.