What legitimate task and unacceptable consequence define the action boundary?
Turn one consequential AI boundary into a decision-ready security record.
Select an AI estate, an agent action path, or an exact release candidate. Define the boundary, examine the relevant exposure and failure paths, and structure the evidence required for an adoption, runtime-control, or release decision.
- 01One bounded decision
- 02Written authority before testing
- 03Evidence tied to material change
- 01BOUNDDecision, owner, target, authority
- 02MAPContext, reachability, controls
- 03EXERCISEAuthorized failure paths
- 04DECIDEEvidence, conditions, limits
- 05VERIFYReplay, expiry, next action
One assessment system. Three precise entry boundaries.
The selected path changes the scope, method, evidence structure, participating teams, limitations, and next conversation, not only the card title.
Agent Action Control Assessment
Evaluate the complete action context, from instruction and identity to retrieved data, delegated authority, tool arguments, target, control response, and replay evidence.
- Boundary
- One consequential agent-to-tool action path
- Method
- Bound → Map → Exercise → Decide → Verify
- Output
- Action-boundary control and replay record
- Owner
- Security engineering and application owner
- Next step
- Qualify fit, then agree written scope
An agent can retrieve sensitive context, invoke tools, delegate authority, write, execute, or change another system.
Follow the evidence from starting object to decision record.
- 01InstructionUser and task intent
- 02IdentityActor and delegated authority
- 03Agent contextMemory and retrieval
- 04Tool requestArguments and target
- 05Control pointAllow, restrict, escalate, stop
- 06Replay recordTrace, response, verification
Illustrative assessment boundary. It is not a customer environment, product capture, supported-technology statement, or measured result.
Coverage begins with decision questions, not unsupported compatibility claims.
Which actor and delegated authority are attached to the proposed action?
Can memory, retrieval, or another agent alter the action path?
Do the tool, arguments, and target remain inside approved purpose and scope?
Should the action be allowed, restricted, transformed, escalated, or stopped?
Can the exact unsafe case be reproduced after the control changes?
These are assessment questions. Actual target, method, depth, access, and evidence are established through written scope.
Specific enough to evaluate. Narrow enough to authorize.
- The initiating user or system, agent identity, delegated authority, prompt or instruction context, and intended business purpose
- Retrieved content, memory, MCP server or tool, arguments, target resource, and proposed consequence
- Existing policy point, allowed response, denied response, escalation path, trace, and replay evidence
- Adversarial prompts, indirect injection, tool manipulation, excessive-agency, or data-exfiltration cases agreed in the rules of engagement
- Identity and permission analysis across adjacent agents when delegation changes the reachable consequence
- Non-production control deployment or replay when a remediation must be verified
- Unbounded production testing, destructive action, persistence, denial of service, or targets outside written authorization
- Claims that one tested path represents every user, agent, model, tool, or future system state
- Secrets, exploit details, customer records, or access material submitted through the public website
Bound. Map. Exercise. Decide. Verify.
- 01Bound
Name the exact action, actor, agent, tool, target, environment, authorized tests, safeguards, and stop conditions.
Action-path scope record - 02Map
Assemble identity, instruction, purpose, memory, retrieval, tool, arguments, target, existing controls, and expected safe path.
Context and control map - 03Exercise
Run only the agreed failure cases and capture the complete trace, observed consequence, control response, and residual exposure.
Reproducible action traces - 04Decide
Identify where the action should be allowed, constrained, escalated, transformed, or stopped and what evidence supports that response.
Control decision record - 05Verify
Replay the exact case against the selected control and record whether the unsafe path closes without breaking the legitimate task.
Replay and retest record
Inspect the record before discussing the engagement.
Runtime action and replay record
A reconstructable record of the instruction, actor, delegated authority, context, tool request, target, control response, evidence, and replay state.
- 01Action boundaryRecorded
- Support agent → customer-record update tool
- 02Authority contextReview
- Service identity present; requested target exceeds reviewed customer scope
- 03Observed failure pathRecorded
- Retrieved instruction alters the target and expands requested fields
- 04Selected controlReview
- Bind target to authorized customer; escalate broader update
- 05Replay stateOpen
- Exact regression case remains required after control change
The record remains tied to its named system, versions, environment, evidence sources, owner, limitations, observation period, and change trigger.
Authority and evidence remain shared, explicit, and bounded.
The assessment begins only when the people responsible for the system, evaluation, safeguards, and resulting decision understand their part.
Authorized system owner
- Confirm lawful authority and accountable ownership
- Describe the intended business purpose and unacceptable consequence
- Approve the environment, access route, evidence handling, and participant list
Cosmipher
- Propose the minimum useful assessment boundary
- Define the method, evidence structure, safeguards, and limitations
- Keep observations tied to the agreed system, versions, environment, and period
Agreed jointly
- Rules of engagement and excluded systems
- Stop conditions, escalation contacts, schedule, and communication
- Decision owner, remediation handoff, evidence validity, retest, and closure
Prepare the decision before preparing the data.
A detailed technical transfer is not the first step. Start with authority, ownership, the decision to be made, the unacceptable consequence, and the smallest useful boundary.
- A named owner can authorize the target.
- One adoption, action, or release decision is explicit.
- A safe environment or controlled method can be discussed.
- The relevant technical and business owners can participate.
- Detailed evidence can move through an approved channel after scope.
- Stop conditions and escalation contacts can be agreed before testing.
A decision is only as current as the system it describes.
The record must expose what was examined, what remained outside scope, and which change requires replay or renewal.
- 01Baseline bound
System, versions, environment, evidence sources, owner, scope, and limitations are recorded.
- 02Evidence assembled
Observed exposure, exercised cases, control state, remediation, and unresolved questions are joined.
- 03Decision recorded
The accountable owner can interpret the result within the stated boundary and validity period.
- 04Material change
A model, prompt, retrieval, identity, permission, tool, artifact, data, policy, purpose, or environment changes.
- 05Renew or invalidate
Affected evidence is replayed, renewed, superseded, or explicitly marked no longer sufficient.
An assessment does not certify that an AI system is universally secure, safe, compliant, or free from future failure. It supports a decision inside the named scope, versions, environment, methods, evidence sources, exclusions, and observation period.
Know what is published, scoped, and demonstrated by evidence.
The public method defines how the decision is structured. The engagement determines the exact boundary. Only observed evidence can support a finding or outcome.
Published method
Decision paths, boundary model, responsibility, assessment stages, record structure, limitations, and evidence-validity logic.
Defined during scope
Target, environment, access, testing depth, safeguards, stop conditions, participants, evidence handling, output format, and retest treatment.
Supported by engagement evidence
Observed findings, exercised coverage, control response, performance, remediation state, replay result, and the decision the evidence can support.
Resolve the operating questions before detailed scoping.
These boundaries keep the first conversation useful without turning a public request into an unsafe technical handoff.
01What should we share through the website or first email?+
Share only the decision, system type, lifecycle stage, responsible team, and high-level non-sensitive context. Do not send credentials, source code, production data, customer records, vulnerability details, exploit material, or confidential architecture through the public channel.
02Does contacting Cosmipher authorize access or testing?+
No. A message or form submission begins qualification only. Access, scanning, installation, connection, testing, or exploitation requires the authorized owner, written scope, an approved environment, rules of engagement, safeguards, contacts, and stop conditions.
03Can an assessment begin outside production?+
The environment is selected during scoping according to the decision, representative behavior required, available safeguards, and the system owner's authorization. The website does not imply that production access is necessary or accepted.
04Which teams should participate?+
The accountable system or release owner should participate. Depending on the path, the working group can include security, AI platform, application engineering, AppSec, IAM, SecOps, MLOps, data owners, model risk, privacy, GRC, and the business control owner.
05How is this different from a product demo?+
A demo explains product workflows using approved or simulated material. An assessment considers one agreed customer boundary and can include examination or testing only after separate written authorization. A demo request never authorizes assessment activity.
06Does an assessment certify that the system is secure or compliant?+
No. The resulting evidence is limited to the agreed target, environment, versions, methods, evidence sources, exclusions, and time period. Framework mapping can organize technical evidence but does not create legal advice, certification, or a universal guarantee.
07What happens when the AI system changes?+
The decision record identifies material-change triggers. A relevant change can invalidate all or part of the earlier evidence and require targeted replay, evidence renewal, or a newly bounded assessment.
08How are detailed evidence and access information handled?+
The first public contact does not collect them. If detailed material is required, the parties first agree the approved channel, purpose, access method, participants, handling restrictions, retention, and deletion expectations as part of scope.
Bring the decision. Keep sensitive material out of the first message.
Share the assessment path, system type, lifecycle stage, responsible team, and high-level decision. Access details and technical evidence belong in an approved channel after authority, scope, safeguards, and handling conditions are agreed.
- No credentials, source code, datasets, vulnerability details, or customer records
- No production access implied or requested
- No testing before written authorization and rules of engagement
- 01Initial context
Share the path, decision, owner, lifecycle stage, and non-sensitive system context.
- 02Mutual fit
Confirm that the decision and boundary match the assessment method and responsible teams.
- 03Written scope
Agree target, authority, exclusions, safeguards, stop conditions, evidence handling, and decision owner.
- 04Authorized start
Assessment activity begins only after the approved scope and rules of engagement are in place.
Discuss one assessment boundary with Cosmipher.
The prepared email includes only the decision fields needed for an initial, non-sensitive scoping conversation.
Start the scoping email- Contact
- Services@cosmipher.com
- Request state
- Qualification only
- Testing authority
- Not granted
Use Contact for platform evaluation, procurement and security review, research, partnerships, or another business question.