Home

Wiki

Review

Docs

PR #8961

  • 1New DeveloperInstructions Builder

10 files+237−28

user_instructions.rscodex-rs/core/src

BUILD.bazelcodex-rs/protocol

models.rscodex-rs/protocol/src

never.mdcodex-rs/protocol/src/prompts/permissions/approval_policy

on_failure.mdcodex-rs/protocol/src/prompts/permissions/approval_policy

+5 more files

Read explanation

The DeveloperInstructions struct moved from core/src/user_instructions.rs to protocol/src/models.rs and gained a new factory method that constructs permissions messages from runtime configuration.

models.rs:189-231

rust

impl DeveloperInstructions {    pub fn new<T: Into<String>>(text: T) -> Self    pub fn into_text(self) -> String    pub fn concat(self, other: impl Into<Self>) -> Self    /// Main entry point: builds permissions    /// message from current policy settings    pub fn from_policy(        sandbox_policy: &SandboxPolicy,        approval_policy: AskForApproval,        cwd: &Path,    ) -> Self {        // Derive network_access from sandbox_policy        // Derive sandbox_mode and writable_roots        // Combine into structured message    }}

The message content comes from template files that get include_str!'d at compile time:

prompts/permissions/├── approval_policy/│   ├── never.md│   ├── on_failure.md│   ├── on_request.md│   └── unless_trusted.md└── sandbox_mode/    ├── danger_full_access.md    ├── read_only.md    └── workspace_write.md

Each template contains the prose for that mode. For example, workspace_write.mdworkspace_write.md:1:

markdown

Filesystem sandboxing defines which filescan be read or written. `sandbox_mode` is`workspace-write`: The sandbox permitsreading files, and editing files in `cwd`and `writable_roots`. Editing files inother directories requires approval.Network access is {network_access}.

The {network_access} placeholder is replaced with enabled or restricted at runtime models.rs:272-281.

  • 2Permissions Message Injection in Session

1 file+51−7

codex.rscodex-rs/core/src

Read explanation

The permissions message is injected at three points in the session lifecycle.

At session start — added to build_initial_context(): codex.rs:1364-1374

rust

fn build_initial_context(&self, ctx: &TurnContext)    -> Vec<ResponseItem>{    let mut items = Vec::with_capacity(4); // was 3        // NEW: permissions message first    items.push(        DeveloperInstructions::from_policy(            &ctx.sandbox_policy,            ctx.approval_policy,            &ctx.cwd,        ).into()    );        // Then: developer_instructions,    //       user_instructions, environment    ...}

When policies change mid-session — new method build_permissions_update_item(): codex.rs:1012-1032

rust

fn build_permissions_update_item(    &self,    previous: Option<&Arc<TurnContext>>,    next: &TurnContext,) -> Option<ResponseItem> {    let prev = previous?;    if prev.sandbox_policy == next.sandbox_policy        && prev.approval_policy == next.approval_policy    {        return None;  // no change, no message    }    Some(DeveloperInstructions::from_policy(...).into())}

This is called during user input handling codex.rs:1895-1901 alongside the existing environment update check.

On resume/fork — initial context appended after reconstructed history: codex.rs:855-859

rust

// After reconstructing history from rollout:let initial_context =    self.build_initial_context(&turn_context);self.record_conversation_items(    &turn_context, &initial_context).await;
  • 3Simplified EnvironmentContext

1 file+20−223

environment_context.rscodex-rs/core/src

Read explanation

Since sandbox/approval info moved to the permissions message, EnvironmentContext was slimmed down significantly.

environment_context.rs:17-25

diff

    pub struct EnvironmentContext {      pub cwd: Option<PathBuf>,-     pub approval_policy: Option<AskForApproval>,-     pub sandbox_mode: Option<SandboxMode>,-     pub network_access: Option<NetworkAccess>,-     pub writable_roots: Option<Vec<AbsolutePathBuf>>,      pub shell: Shell,  }

The constructor simplified from a complex match on SandboxPolicy to: environment_context.rs:19-20

rust

pub fn new(cwd: Option<PathBuf>, shell: Shell) -> Self {    Self { cwd, shell }}

The equals_except_shell() comparison now only checks cwdenvironment_context.rs:27-34, and the XML serialization no longer emits approval/sandbox/network tags environment_context.rs:51-70.

  • 4Removed Static Prompt Sections

7 files−255

gpt-5.1-codex-max_prompt.mdcodex-rs/core

gpt-5.2-codex_prompt.mdcodex-rs/core

gpt_5_1_prompt.mdcodex-rs/core

gpt_5_2_prompt.mdcodex-rs/core

gpt_5_codex_prompt.mdcodex-rs/core

+2 more files

Read explanation

The "Codex CLI harness, sandboxing, and approvals" section (~35 lines) was removed from all prompt markdown files since this content is now generated dynamically:

2 files+449

mod.rscodex-rs/core/tests/suite

permissions_messages.rscodex-rs/core/tests/suite

Read explanation

A comprehensive test file validates the new permissions messaging behavior permissions_messages.rs:1-448:

  • permissions_message_sent_once_on_start — verifies single emission at session start

    • permissions_message_added_on_override_change — verifies re-emission when policy changes
    • permissions_message_not_added_when_no_change — verifies no duplicate when policy unchanged
    • resume_replays_permissions_messages — verifies history replay on resume
    • resume_and_fork_append_permissions_messages — verifies fresh context appended on resume/fork
    • permissions_message_includes_writable_roots — verifies writable roots formatting
  • 6Test Adjustments for New Message Structure

9 files+332−142

send_message.rscodex-rs/app-server/tests/suite

truncation.rscodex-rs/core/src/rollout

thread_manager.rscodex-rs/core/src

client.rscodex-rs/core/tests/suite

compact.rscodex-rs/core/tests/suite

+4 more files

Read explanation

Tests throughout the codebase were updated to account for the new permissions message at input[0]. Most changes are mechanical index shifts (e.g., input[0] → input[1]) and additions of permissions_message to expected JSON structures.

Key patterns:

  • Request body assertions now check for permissions message first client.rs:652-659
    • Resume/fork tests verify message ordering client.rs:345-349
    • Compact tests filter out permissions messages when comparing compact.rs:607-614
    • Fork tests account for initial context appended after truncation fork_thread.rs:141-143

Merged

openai/codex #8961

Assemble sandbox/approval/network prompts dynamically

AIaibrahim-oai

main dynamic/permissions/instructions

30 files

+1089−655

DescriptionDiscussion41Commits48

Devin's AI analysis

This PR moves sandbox/approval/network permission instructions from static system prompt files into dynamically generated developer messages. Previously, these instructions were baked into markdown files and sent identically to every session. Now they're constructed at runtime based on actual configuration.

Key change: A new DeveloperInstructions::from_policy() method models.rs:204-231 builds a developer-role message from the current SandboxPolicy and AskForApproval settings. This message is:

  • Injected at session start via build_initial_context()codex.rs:1367-1374
  • Re-injected when policies change mid-session via build_permissions_update_item()codex.rs:1012-1032
  • Appended after reconstructed history on resume/fork codex.rs:855-859

Consequence: EnvironmentContext no longer carries approval_policy, sandbox_mode, network_access, or writable_roots fields—these moved to the permissions message. The struct now only holds cwd and shell.

Request structure change:

Before:  input[0]: user instructions (AGENTS.md)  input[1]: environment context  input[2]: user messageAfter:  input[0]: permissions (developer role)  input[1]: user instructions (AGENTS.md)  input[2]: environment context  input[3]: user message
  • Add a single builder for developer permissions messaging that accepts SandboxPolicy and approval policy. This builder now drives the developer “permissions” message that’s injected at session start and any time sandbox/approval settings change.
  • Trim EnvironmentContext to only include cwd, writable roots, and shell; removed sandbox/approval/network duplication and adjusted XML serialization and tests accordingly.

Follow-up: adding a config value to replace the developer permissions message for custom sandboxes.

Filter

AI

aibrahim-oai112d

@codex review

CH

chatgpt-codex-connector

commented

112d

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:a40e5a755a

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

models.rsOutdatedResolved

AI

aibrahim-oai112d

@codex review this and make sure tests changes didn't invalidate tests original intentions

CH

chatgpt-codex-connector

commented

112d

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:2240249141

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

prompt_caching.rsOutdatedResolved

AI

aibrahim-oai112d

@codex review this

CH

chatgpt-codex-connector

commented

112d

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:fc5c68019b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

on_request.mdOutdatedResolved

AI

aibrahim-oai112d

@codex review

CH

chatgpt-codex-connector112d

Codex Review: Didn't find any major issues. Swish!

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

PApakrym-oai

commented

112d

models.rsL175

PA

pakrym-oai112d

a bit strange to see all of this formatting stuff in the protocol

AI

aibrahim-oai112d

Yeah, just so that they can be methods on the structs. I can move them with the structs to core.

DY

dylan-hurd-oai109d

Agree w/ pavel. Is the plan to start using DeveloperInstructions as a serialized struct?

AI

aibrahim-oai109d

no I can move it

PApakrym-oai

commented

112d

models.rsOutdated

PA

pakrym-oai112d

A lot of methods. Do we have this many callsites that we need all the overrides?

AI

aibrahim-oai112d

Not really, we can have one method but just wanted to divide responsibility so we don't have a super long function with complicated logic.

PA

pakrym-oai112d

You can return strings from helper method and make them private.

PApakrym-oai

commented

112d

models.rsOutdated

PA

pakrym-oai112d

Why are these values encoded here? is there an existing place where we convert SandboxPolicy to SandboxMode/network?

PApakrym-oai

commented

112d

models.rsL289Resolved

PApakrym-oai

commented

112d

thread_manager.rsL405

PA

pakrym-oai112d

is this expected?

👍1

AI

aibrahim-oai109d

yes we added a new message

PApakrym-oai

commented

112d

environment_context.rsL14

PA

pakrym-oai112d

we left so little here I wonder whether it matters.

AI

aibrahim-oai112d

I want to move it to dev slowly per the plan.

AI

aibrahim-oai111d

@codex do we append a new permission message on resume and fork next to the older messages? If yes, add an integration test. If not, change behavior and add integration test.

CH

chatgpt-codex-connector111d

Summary

Testing

  • ✅ /root/.cargo/bin/just fmt
  • ✅ /root/.cargo/bin/just fix -p codex-core
  • ❌ cargo test -p codex-core (failed: default_client::tests::test_create_client_sets_default_headers; exec::tests::kill_child_process_group_kills_grandchildren_on_timeout; shell_snapshot::tests::linux_sh_snapshot_includes_sections; models_manager::manager::tests::refresh_available_models_{drops_removed_remote_models,refetches_when_cache_stale,sorts_and_marks_default,uses_cache_when_fresh}; unified_exec::tests::{reusing_completed_process_returns_unknown_process,multi_unified_exec_sessions})

View task →

DY

dylan-hurd-oai

commented

109d

Some small comments but mostly LGTM!

codex.rsL1032

DY

dylan-hurd-oai109d

(non-blocking) Should we move this out of codex.rs? Either into environment_context.rs or a new file

AI

aibrahim-oai109d

it's on turn context unfortunately

codex.rsL2982Resolved

1,744 lines left

1New DeveloperInstructions Builder

0 / 10

user_instructions.rscodex-rs/core/src

−28

Mark as viewed

5 linesAll 77 lines5 lines

5 linesAll 77 lines5 lines

1

Move78–104

78

#[derive(Debug, Clone, Serialize, Deserialize, PartialEq)]

79

#[serde(rename = "developer_instructions", rename_all = "snake_case")]

80

pub(crate) struct DeveloperInstructions {

81

text: String,

82

}

83

84

impl DeveloperInstructions {

85

pub fn new<T: Into<String>>(text: T) -> Self {

86

Self { text: text.into() }

87

}

88

89

pub fn into_text(self) -> String {

90

self.text

91

}

92

}

93

94

impl From<DeveloperInstructions> for ResponseItem {

95

fn from(di: DeveloperInstructions) -> Self {

96

ResponseItem::Message {

97

id: None,

98

role: "developer".to_string(),

99

content: vec![ContentItem::InputText {

100

text: di.into_text(),

101

}],

102

}

103

}

104

}

105

5 linesAll 88 lines5 lines

5 linesAll 88 lines5 lines

BUILD.bazelcodex-rs/protocol

+1

Mark as viewed

1

load("//:defs.bzl", "codex_rust_crate")

1

load("//:defs.bzl", "codex_rust_crate")

2

``

2

``

3

codex_rust_crate(

3

codex_rust_crate(

4

name = "protocol",

4

name = "protocol",

5

crate_name = "codex_protocol",

5

crate_name = "codex_protocol",

6

compile_data = glob(["src/prompts/permissions/**/*.md"]),

6

)

7

)

models.rscodex-rs/protocol/src

+218

Mark as viewed

1

use std::collections::HashMap;

1

use std::collections::HashMap;

2

use std::path::Path;

5 linesAll 10 lines5 lines

5 linesAll 10 lines5 lines

13

use crate::config_types::SandboxMode;

14

use crate::protocol::AskForApproval;

15

use crate::protocol::NetworkAccess;

16

use crate::protocol::SandboxPolicy;

17

use crate::protocol::WritableRoot;

5 linesAll 149 lines5 lines

5 linesAll 149 lines5 lines

1

Move167–294

167

/// Developer-provided guidance that is injected into a turn as a developer role

168

/// message.

169

#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, JsonSchema, TS)]

170

#[serde(rename = "developer_instructions", rename_all = "snake_case")]

171

pub struct DeveloperInstructions {

172

text: String,

173

}

174

PA4

CommentR175

175

const APPROVAL_POLICY_NEVER: &str = include_str!("prompts/permissions/approval_policy/never.md");

176

const APPROVAL_POLICY_UNLESS_TRUSTED: &str =

177

include_str!("prompts/permissions/approval_policy/unless_trusted.md");

178

const APPROVAL_POLICY_ON_FAILURE: &str =

179

include_str!("prompts/permissions/approval_policy/on_failure.md");

180

const APPROVAL_POLICY_ON_REQUEST: &str =

181

include_str!("prompts/permissions/approval_policy/on_request.md");

182

183

const SANDBOX_MODE_DANGER_FULL_ACCESS: &str =

184

include_str!("prompts/permissions/sandbox_mode/danger_full_access.md");

185

const SANDBOX_MODE_WORKSPACE_WRITE: &str =

186

include_str!("prompts/permissions/sandbox_mode/workspace_write.md");

187

const SANDBOX_MODE_READ_ONLY: &str = include_str!("prompts/permissions/sandbox_mode/read_only.md");

188

189

impl DeveloperInstructions {

190

pub fn new<T: Into<String>>(text: T) -> Self {

191

Self { text: text.into() }

192

}

193

194

pub fn into_text(self) -> String {

195

self.text

196

}

197

198

pub fn concat(self, other: impl Into<DeveloperInstructions>) -> Self {

199

let mut text = self.text;

200

text.push_str(&other.into().text);

201

Self { text }

202

}

203

204

pub fn from_policy(

205

sandbox_policy: &SandboxPolicy,

206

approval_policy: AskForApproval,

207

cwd: &Path,

208

) -> Self {

209

let network_access = if sandbox_policy.has_full_network_access() {

210

NetworkAccess::Enabled

211

} else {

212

NetworkAccess::Restricted

213

};

214

215

let (sandbox_mode, writable_roots) = match sandbox_policy {

216

SandboxPolicy::DangerFullAccess => (SandboxMode::DangerFullAccess, None),

217

SandboxPolicy::ReadOnly => (SandboxMode::ReadOnly, None),

218

SandboxPolicy::ExternalSandbox { .. } => (SandboxMode::DangerFullAccess, None),

219

SandboxPolicy::WorkspaceWrite { .. } => {

220

let roots = sandbox_policy.get_writable_roots_with_cwd(cwd);

221

(SandboxMode::WorkspaceWrite, Some(roots))

222

}

223

};

224

225

DeveloperInstructions::from_permissions_with_network(

226

sandbox_mode,

227

network_access,

228

approval_policy,

229

writable_roots,

230

)

231

}

232

233

fn from_permissions_with_network(

234

sandbox_mode: SandboxMode,

235

network_access: NetworkAccess,

236

approval_policy: AskForApproval,

237

writable_roots: Option<Vec<WritableRoot>>,

238

) -> Self {

239

let start_tag = DeveloperInstructions::new("<permissions instructions>");

240

let end_tag = DeveloperInstructions::new("</permissions instructions>");

241

start_tag

242

.concat(DeveloperInstructions::sandbox_text(

243

sandbox_mode,

244

network_access,

245

))

246

.concat(DeveloperInstructions::from(approval_policy))

247

.concat(DeveloperInstructions::from_writable_roots(writable_roots))

248

.concat(end_tag)

249

}

250

251

fn from_writable_roots(writable_roots: Option<Vec<WritableRoot>>) -> Self {

252

let Some(roots) = writable_roots else {

253

return DeveloperInstructions::new("");

254

};

255

256

if roots.is_empty() {

257

return DeveloperInstructions::new("");

258

}

259

260

let roots_list: Vec<String> = roots

261

.iter()

262

.map(|r| format!("`{}`", r.root.to_string_lossy()))

263

.collect();

264

let text = if roots_list.len() == 1 {

265

format!(" The writable root is {}.", roots_list[0])

266

} else {

267

format!(" The writable roots are {}.", roots_list.join(", "))

268

};

269

DeveloperInstructions::new(text)

270

}

271

InformationalR272-281

272

fn sandbox_text(mode: SandboxMode, network_access: NetworkAccess) -> DeveloperInstructions {

273

let template = match mode {

274

SandboxMode::DangerFullAccess => SANDBOX_MODE_DANGER_FULL_ACCESS.trim_end(),

275

SandboxMode::WorkspaceWrite => SANDBOX_MODE_WORKSPACE_WRITE.trim_end(),

276

SandboxMode::ReadOnly => SANDBOX_MODE_READ_ONLY.trim_end(),

277

};

278

let text = template.replace("{network_access}", &network_access.to_string());

279

280

DeveloperInstructions::new(text)

281

}

282

}

283

284

impl From<DeveloperInstructions> for ResponseItem {

285

fn from(di: DeveloperInstructions) -> Self {

286

ResponseItem::Message {

287

id: None,

288

role: "developer".to_string(),

PA

CommentR289

Resolved

289

content: vec![ContentItem::InputText {

290

text: di.into_text(),

291

}],

292

}

293

}

294

}

295

296

impl From<SandboxMode> for DeveloperInstructions {

297

fn from(mode: SandboxMode) -> Self {

298

let network_access = match mode {

299

SandboxMode::DangerFullAccess => NetworkAccess::Enabled,

300

SandboxMode::WorkspaceWrite | SandboxMode::ReadOnly => NetworkAccess::Restricted,

301

};

302

303

DeveloperInstructions::sandbox_text(mode, network_access)

304

}

305

}

306

307

impl From<AskForApproval> for DeveloperInstructions {

308

fn from(mode: AskForApproval) -> Self {

309

let text = match mode {

310

AskForApproval::Never => APPROVAL_POLICY_NEVER.trim_end(),

311

AskForApproval::UnlessTrusted => APPROVAL_POLICY_UNLESS_TRUSTED.trim_end(),

312

AskForApproval::OnFailure => APPROVAL_POLICY_ON_FAILURE.trim_end(),

313

AskForApproval::OnRequest => APPROVAL_POLICY_ON_REQUEST.trim_end(),

314

};

315

316

DeveloperInstructions::new(text)

317

}

318

}

319

5 linesAll 389 lines5 lines

5 linesAll 389 lines5 lines

550

#[cfg(test)]

709

#[cfg(test)]

551

mod tests {

710

mod tests {

552

use super::*;

711

use super::*;

712

use crate::config_types::SandboxMode;

713

use crate::protocol::AskForApproval;

All 4 lines

All 4 lines

718

use std::path::PathBuf;

557

use tempfile::tempdir;

719

use tempfile::tempdir;

558

720

721

#[test]

722

fn converts_sandbox_mode_into_developer_instructions() {

723

assert_eq!(

724

DeveloperInstructions::from(SandboxMode::WorkspaceWrite),

725

DeveloperInstructions::new(

726

"Filesystem sandboxing defines which files can be read or written. `sandbox_mode` is `workspace-write`: The sandbox permits reading files, and editing files in `cwd` and `writable_roots`. Editing files in other directories requires approval. Network access is restricted."

727

)

728

);

729

730

assert_eq!(

731

DeveloperInstructions::from(SandboxMode::ReadOnly),

732

DeveloperInstructions::new(

733

"Filesystem sandboxing defines which files can be read or written. `sandbox_mode` is `read-only`: The sandbox only permits reading files. Network access is restricted."

734

)

735

);

736

}

737

738

#[test]

739

fn builds_permissions_with_network_access_override() {

740

let instructions = DeveloperInstructions::from_permissions_with_network(

741

SandboxMode::WorkspaceWrite,

742

NetworkAccess::Enabled,

743

AskForApproval::OnRequest,

744

None,

745

);

746

747

let text = instructions.into_text();

748

assert!(

749

text.contains("Network access is enabled."),

750

"expected network access to be enabled in message"

751

);

752

assert!(

753

text.contains("`approval_policy` is `on-request`"),

754

"expected approval guidance to be included"

755

);

756

}

757

758

#[test]

759

fn builds_permissions_from_policy() {

760

let policy = SandboxPolicy::WorkspaceWrite {

761

writable_roots: vec![],

762

network_access: true,

763

exclude_tmpdir_env_var: false,

764

exclude_slash_tmp: false,

765

};

766

767

let instructions = DeveloperInstructions::from_policy(

768

&policy,

769

AskForApproval::UnlessTrusted,

770

&PathBuf::from("/tmp"),

771

);

772

let text = instructions.into_text();

773

assert!(text.contains("Network access is enabled."));

774

assert!(text.contains("`approval_policy` is `unless-trusted`"));

775

}

776

5 linesAll 311 lines5 lines

5 linesAll 311 lines5 lines

870

}

1088

}

never.mdcodex-rs/protocol/src/prompts/permissions/approval_policy

+1

AddedMark as viewed

1

Approvals are your mechanism to get user consent to run shell commands without the sandbox. `approval_policy` is `never`: This is a non-interactive mode where you may NEVER ask the user for approval to run commands. Instead, you must always persist and work around constraints to solve the task for the user. You MUST do your utmost best to finish the task and validate your work before yielding. If this mode is paired with `danger-full-access`, take advantage of it to deliver the best outcome for the user. Further, in this mode, your default testing philosophy is overridden: Even if you don't see local patterns for testing, you may add tests and scripts to validate your work. Just remove them before yielding.

on_failure.mdcodex-rs/protocol/src/prompts/permissions/approval_policy

+1

AddedMark as viewed

1

Approvals are your mechanism to get user consent to run shell commands without the sandbox. `approval_policy` is `on-failure`: The harness will allow all commands to run in the sandbox (if enabled), and failures will be escalated to the user for approval to run again without the sandbox.

on_request.mdcodex-rs/protocol/src/prompts/permissions/approval_policy

+12

AddedMark as viewed

1

Move1–12

1

Approvals are your mechanism to get user consent to run shell commands without the sandbox. `approval_policy` is `on-request`: Commands will be run in the sandbox by default, and you can specify in your tool call if you want to escalate a command to run without sandboxing. If the completing the task requires escalated permissions, Do not let these settings or the sandbox deter you from attempting to accomplish the user's task.

2

3

Here are scenarios where you'll need to request approval:

4

- You need to run a command that writes to a directory that requires it (e.g. running tests that write to /var)

5

- You need to run a GUI app (e.g., open/xdg-open/osascript) to open browsers or files.

6

- You are running sandboxed and need to run a command that requires network access (e.g. installing packages)

7

- If you run a command that is important to solving the user's query, but it fails because of sandboxing, rerun the command with approval. ALWAYS proceed to use the `sandbox_permissions` and `justification` parameters - do not message the user before requesting approval for the command.

8

- You are about to take a potentially destructive action such as an `rm` or `git reset` that the user did not explicitly ask for.

9

10

When requesting approval to execute a command that will require escalated privileges:

11

`` - Provide the sandbox_permissions parameter with the value `"require_escalated"```

12

- Include a short, 1 sentence explanation for why you need escalated permissions in the justification parameter

unless_trusted.mdcodex-rs/protocol/src/prompts/permissions/approval_policy

+1

AddedMark as viewed

1

Approvals are your mechanism to get user consent to run shell commands without the sandbox. `approval_policy` is `unless-trusted`: The harness will escalate most commands for user approval, apart from a limited allowlist of safe "read" commands.

danger_full_access.mdcodex-rs/protocol/src/prompts/permissions/sandbox_mode

+1

AddedMark as viewed

1

Filesystem sandboxing defines which files can be read or written. `sandbox_mode` is `danger-full-access`: No filesystem sandboxing - all commands are permitted. Network access is {network_access}.

read_only.mdcodex-rs/protocol/src/prompts/permissions/sandbox_mode

+1

AddedMark as viewed

1

Filesystem sandboxing defines which files can be read or written. `sandbox_mode` is `read-only`: The sandbox only permits reading files. Network access is {network_access}.

workspace_write.mdcodex-rs/protocol/src/prompts/permissions/sandbox_mode

+1

AddedMark as viewed

1

Move1–1

1

Filesystem sandboxing defines which files can be read or written. `sandbox_mode` is `workspace-write`: The sandbox permits reading files, and editing files in `cwd` and `writable_roots`. Editing files in other directories requires approval. Network access is {network_access}.

2Permissions Message Injection in Session

0 / 1

codex.rscodex-rs/core/src

+51−7

Mark as viewed

5 linesAll 149 lines5 lines

5 linesAll 149 lines5 lines

150

use crate::unified_exec::UnifiedExecProcessManager;

150

use crate::unified_exec::UnifiedExecProcessManager;

151

use crate::user_instructions::DeveloperInstructions;

152

use crate::user_instructions::UserInstructions;

151

use crate::user_instructions::UserInstructions;

All 4 lines

All 4 lines

157

use codex_protocol::config_types::ReasoningSummary as ReasoningSummaryConfig;

156

use codex_protocol::config_types::ReasoningSummary as ReasoningSummaryConfig;

158

use codex_protocol::models::ContentItem;

157

use codex_protocol::models::ContentItem;

158

use codex_protocol::models::DeveloperInstructions;

5 linesAll 327 lines5 lines

5 linesAll 327 lines5 lines

486

impl Session {

486

impl Session {

5 linesAll 307 lines5 lines

5 linesAll 307 lines5 lines

794

async fn record_initial_history(&self, conversation_history: InitialHistory) {

794

async fn record_initial_history(&self, conversation_history: InitialHistory) {

795

let turn_context = self.new_default_turn().await;

795

let turn_context = self.new_default_turn().await;

796

match conversation_history {

796

match conversation_history {

All 7 lines

All 7 lines

804

InitialHistory::Resumed(_) | InitialHistory::Forked(_) => {

804

InitialHistory::Resumed(_) | InitialHistory::Forked(_) => {

5 linesAll 46 lines5 lines

5 linesAll 46 lines5 lines

851

// If persisting, persist all rollout items as-is (recorder filters)

851

// If persisting, persist all rollout items as-is (recorder filters)

852

if persist && !rollout_items.is_empty() {

852

if persist && !rollout_items.is_empty() {

853

self.persist_rollout_items(&rollout_items).await;

853

self.persist_rollout_items(&rollout_items).await;

854

}

854

}

855

InformationalR856-859

856

// Append the current session's initial context after the reconstructed history.

857

let initial_context = self.build_initial_context(&turn_context);

858

self.record_conversation_items(&turn_context, &initial_context)

859

.await;

855

// Flush after seeding history and any persisted rollout copy.

860

// Flush after seeding history and any persisted rollout copy.

856

self.flush_rollout().await;

861

self.flush_rollout().await;

857

}

862

}

858

}

863

}

859

}

864

}

5 linesAll 147 lines5 lines

5 linesAll 147 lines5 lines

1

Copy + paste1012–1032

1012

fn build_permissions_update_item(

1013

&self,

1014

previous: Option<&Arc<TurnContext>>,

1015

next: &TurnContext,

1016

) -> Option<ResponseItem> {

1

Potential BugR1017-1022

1017

let prev = previous?;

1018

if prev.sandbox_policy == next.sandbox_policy

1019

&& prev.approval_policy == next.approval_policy

1020

{

1021

return None;

1022

}

1023

1024

Some(

1025

DeveloperInstructions::from_policy(

1026

&next.sandbox_policy,

1027

next.approval_policy,

1028

&next.cwd,

1029

)

1030

.into(),

1031

)

DY2

CommentR1032

1032

}

1033

5 linesAll 329 lines5 lines

5 linesAll 329 lines5 lines

1336

1363

1337

pub(crate) fn build_initial_context(&self, turn_context: &TurnContext) -> Vec<ResponseItem> {

1364

pub(crate) fn build_initial_context(&self, turn_context: &TurnContext) -> Vec<ResponseItem> {

1338

let mut items = Vec::<ResponseItem>::with_capacity(3);

1365

let mut items = Vec::<ResponseItem>::with_capacity(4);

1339

let shell = self.user_shell();

1366

let shell = self.user_shell();

1367

items.push(

1368

DeveloperInstructions::from_policy(

1369

&turn_context.sandbox_policy,

1370

turn_context.approval_policy,

1371

&turn_context.cwd,

1372

)

1373

.into(),

1374

);

1340

if let Some(developer_instructions) = turn_context.developer_instructions.as_deref() {

1375

if let Some(developer_instructions) = turn_context.developer_instructions.as_deref() {

5 linesAll 10 lines5 lines

5 linesAll 10 lines5 lines

1351

}

1386

}

1352

items.push(ResponseItem::from(EnvironmentContext::new(

1387

items.push(ResponseItem::from(EnvironmentContext::new(

1353

Some(turn_context.cwd.clone()),

1388

Some(turn_context.cwd.clone()),

1354

Some(turn_context.approval_policy),

1355

Some(turn_context.sandbox_policy.clone()),

1356

shell.as_ref().clone(),

1389

shell.as_ref().clone(),

1357

)));

1390

)));

1358

items

1391

items

1359

}

1392

}

5 linesAll 281 lines5 lines

5 linesAll 281 lines5 lines

1641

}

1674

}

5 linesAll 101 lines5 lines

5 linesAll 101 lines5 lines

1743

mod handlers {

1776

mod handlers {

5 linesAll 60 lines5 lines

5 linesAll 60 lines5 lines

1804

pub async fn user_input_or_turn(

1837

pub async fn user_input_or_turn(

1805

sess: &Arc<Session>,

1838

sess: &Arc<Session>,

1806

sub_id: String,

1839

sub_id: String,

1807

op: Op,

1840

op: Op,

1808

previous_context: &mut Option<Arc<TurnContext>>,

1841

previous_context: &mut Option<Arc<TurnContext>>,

1809

) {

1842

) {

5 linesAll 44 lines5 lines

5 linesAll 44 lines5 lines

1854

// Attempt to inject input into current task

1887

// Attempt to inject input into current task

1855

if let Err(items) = sess.inject_input(items).await {

1888

if let Err(items) = sess.inject_input(items).await {

1889

let mut update_items = Vec::new();

1856

if let Some(env_item) =

1890

if let Some(env_item) =

1857

sess.build_environment_update_item(previous_context.as_ref(), &current_context)

1891

sess.build_environment_update_item(previous_context.as_ref(), &current_context)

1858

{

1892

{

1859

sess.record_conversation_items(&current_context, std::slice::from_ref(&env_item))

1893

update_items.push(env_item);

1894

}

1895

if let Some(permissions_item) =

1896

sess.build_permissions_update_item(previous_context.as_ref(), &current_context)

1897

{

1898

update_items.push(permissions_item);

1899

}

1900

if !update_items.is_empty() {

1901

sess.record_conversation_items(&current_context, &update_items)

1860

.await;

1902

.await;

1861

}

1903

}

All 4 lines

All 4 lines

1866

}

1908

}

1867

}

1909

}

5 linesAll 333 lines5 lines

5 linesAll 333 lines5 lines

2201

}

2243

}

5 linesAll 668 lines5 lines

5 linesAll 668 lines5 lines

2870

mod tests {

2912

mod tests {

5 linesAll 56 lines5 lines

5 linesAll 56 lines5 lines

2927

#[tokio::test]

2969

#[tokio::test]

2928

async fn record_initial_history_reconstructs_resumed_transcript() {

2970

async fn record_initial_history_reconstructs_resumed_transcript() {

2929

let (session, turn_context) = make_session_and_context().await;

2971

let (session, turn_context) = make_session_and_context().await;

2930

let (rollout_items, expected) = sample_rollout(&session, &turn_context);

2972

let (rollout_items, mut expected) = sample_rollout(&session, &turn_context);

2931

2973

All 6 lines

All 6 lines

2938

.await;

2980

.await;

2939

2981

DY

CommentR2982

Resolved

2982

expected.extend(session.build_initial_context(&turn_context));

2940

let history = session.state.lock().await.clone_history();

2983

let history = session.state.lock().await.clone_history();

2941

assert_eq!(expected, history.raw_items());

2984

assert_eq!(expected, history.raw_items());

2942

}

2985

}

5 linesAll 78 lines5 lines

5 linesAll 78 lines5 lines

3021

#[tokio::test]

3064

#[tokio::test]

3022

async fn record_initial_history_reconstructs_forked_transcript() {

3065

async fn record_initial_history_reconstructs_forked_transcript() {

3023

let (session, turn_context) = make_session_and_context().await;

3066

let (session, turn_context) = make_session_and_context().await;

3024

let (rollout_items, expected) = sample_rollout(&session, &turn_context);

3067

let (rollout_items, mut expected) = sample_rollout(&session, &turn_context);

3025

3068

3026

session

3069

session

3027

.record_initial_history(InitialHistory::Forked(rollout_items))

3070

.record_initial_history(InitialHistory::Forked(rollout_items))

3028

.await;

3071

.await;

3029

3072

3073

expected.extend(session.build_initial_context(&turn_context));

3030

let history = session.state.lock().await.clone_history();

3074

let history = session.state.lock().await.clone_history();

3031

assert_eq!(expected, history.raw_items());

3075

assert_eq!(expected, history.raw_items());

3032

}

3076

}

5 linesAll 1108 lines5 lines

5 linesAll 1108 lines5 lines

4141

}

4185

}

3Simplified EnvironmentContext

0 / 1

environment_context.rscodex-rs/core/src

+20−223

Mark as viewed

1

use crate::codex::TurnContext;

1

use crate::codex::TurnContext;

2

use crate::protocol::AskForApproval;

3

use crate::protocol::NetworkAccess;

4

use crate::protocol::SandboxPolicy;

5

use crate::shell::Shell;

2

use crate::shell::Shell;

6

use codex_protocol::config_types::SandboxMode;

7

use codex_protocol::models::ContentItem;

3

use codex_protocol::models::ContentItem;

8

use codex_protocol::models::ResponseItem;

4

use codex_protocol::models::ResponseItem;

9

use codex_protocol::protocol::ENVIRONMENT_CONTEXT_CLOSE_TAG;

5

use codex_protocol::protocol::ENVIRONMENT_CONTEXT_CLOSE_TAG;

10

use codex_protocol::protocol::ENVIRONMENT_CONTEXT_OPEN_TAG;

6

use codex_protocol::protocol::ENVIRONMENT_CONTEXT_OPEN_TAG;

11

use codex_utils_absolute_path::AbsolutePathBuf;

12

use serde::Deserialize;

7

use serde::Deserialize;

13

use serde::Serialize;

8

use serde::Serialize;

14

use std::path::PathBuf;

9

use std::path::PathBuf;

15

10

16

#[derive(Debug, Clone, Serialize, Deserialize, PartialEq)]

11

#[derive(Debug, Clone, Serialize, Deserialize, PartialEq)]

InformationalR12-16

17

#[serde(rename = "environment_context", rename_all = "snake_case")]

12

#[serde(rename = "environment_context", rename_all = "snake_case")]

18

pub(crate) struct EnvironmentContext {

13

pub(crate) struct EnvironmentContext {

PA2

CommentR14

19

pub cwd: Option<PathBuf>,

14

pub cwd: Option<PathBuf>,

20

pub approval_policy: Option<AskForApproval>,

21

pub sandbox_mode: Option<SandboxMode>,

22

pub network_access: Option<NetworkAccess>,

23

pub writable_roots: Option<Vec<AbsolutePathBuf>>,

24

pub shell: Shell,

15

pub shell: Shell,

25

}

16

}

26

17

27

impl EnvironmentContext {

18

impl EnvironmentContext {

28

pub fn new(

19

pub fn new(cwd: Option<PathBuf>, shell: Shell) -> Self {

29

cwd: Option<PathBuf>,

20

Self { cwd, shell }

30

approval_policy: Option<AskForApproval>,

31

sandbox_policy: Option<SandboxPolicy>,

32

shell: Shell,

33

) -> Self {

34

Self {

35

cwd,

36

approval_policy,

37

sandbox_mode: match sandbox_policy {

38

Some(SandboxPolicy::DangerFullAccess) => Some(SandboxMode::DangerFullAccess),

39

Some(SandboxPolicy::ReadOnly) => Some(SandboxMode::ReadOnly),

40

Some(SandboxPolicy::ExternalSandbox { .. }) => Some(SandboxMode::DangerFullAccess),

41

Some(SandboxPolicy::WorkspaceWrite { .. }) => Some(SandboxMode::WorkspaceWrite),

42

None => None,

43

},

44

network_access: match sandbox_policy {

45

Some(SandboxPolicy::DangerFullAccess) => Some(NetworkAccess::Enabled),

46

Some(SandboxPolicy::ReadOnly) => Some(NetworkAccess::Restricted),

47

Some(SandboxPolicy::ExternalSandbox { network_access }) => Some(network_access),

48

Some(SandboxPolicy::WorkspaceWrite { network_access, .. }) => {

49

if network_access {

50

Some(NetworkAccess::Enabled)

51

} else {

52

Some(NetworkAccess::Restricted)

53

}

54

}

55

None => None,

56

},

57

writable_roots: match sandbox_policy {

58

Some(SandboxPolicy::WorkspaceWrite { writable_roots, .. }) => {

59

if writable_roots.is_empty() {

60

None

61

} else {

62

Some(writable_roots)

63

}

64

}

65

_ => None,

66

},

67

shell,

68

}

69

}

21

}

70

22

71

/// Compares two environment contexts, ignoring the shell. Useful when

23

/// Compares two environment contexts, ignoring the shell. Useful when

72

/// comparing turn to turn, since the initial environment_context will

24

/// comparing turn to turn, since the initial environment_context will

73

/// include the shell, and then it is not configurable from turn to turn.

25

/// include the shell, and then it is not configurable from turn to turn.

74

pub fn equals_except_shell(&self, other: &EnvironmentContext) -> bool {

26

pub fn equals_except_shell(&self, other: &EnvironmentContext) -> bool {

75

let EnvironmentContext {

27

let EnvironmentContext {

76

cwd,

28

cwd,

77

approval_policy,

78

sandbox_mode,

79

network_access,

80

writable_roots,

81

// should compare all fields except shell

29

// should compare all fields except shell

82

shell: _,

30

shell: _,

83

} = other;

31

} = other;

84

32

85

self.cwd == *cwd

33

self.cwd == *cwd

86

&& self.approval_policy == *approval_policy

87

&& self.sandbox_mode == *sandbox_mode

88

&& self.network_access == *network_access

89

&& self.writable_roots == *writable_roots

90

}

34

}

91

35

92

pub fn diff(before: &TurnContext, after: &TurnContext, shell: &Shell) -> Self {

36

pub fn diff(before: &TurnContext, after: &TurnContext, shell: &Shell) -> Self {

93

let cwd = if before.cwd != after.cwd {

37

let cwd = if before.cwd != after.cwd {

94

Some(after.cwd.clone())

38

Some(after.cwd.clone())

95

} else {

39

} else {

96

None

40

None

97

};

41

};

98

let approval_policy = if before.approval_policy != after.approval_policy {

42

EnvironmentContext::new(cwd, shell.clone())

99

Some(after.approval_policy)

100

} else {

101

None

102

};

103

let sandbox_policy = if before.sandbox_policy != after.sandbox_policy {

104

Some(after.sandbox_policy.clone())

105

} else {

106

None

107

};

108

EnvironmentContext::new(cwd, approval_policy, sandbox_policy, shell.clone())

109

}

43

}

110

44

111

pub fn from_turn_context(turn_context: &TurnContext, shell: &Shell) -> Self {

45

pub fn from_turn_context(turn_context: &TurnContext, shell: &Shell) -> Self {

112

Self::new(

46

Self::new(Some(turn_context.cwd.clone()), shell.clone())

113

Some(turn_context.cwd.clone()),

114

Some(turn_context.approval_policy),

115

Some(turn_context.sandbox_policy.clone()),

116

shell.clone(),

117

)

118

}

47

}

119

}

48

}

120

49

121

impl EnvironmentContext {

50

impl EnvironmentContext {

122

`` /// Serializes the environment context to XML. Libraries like `quick-xml```

51

`` /// Serializes the environment context to XML. Libraries like `quick-xml```

123

/// require custom macros to handle Enums with newtypes, so we just do it

52

/// require custom macros to handle Enums with newtypes, so we just do it

124

/// manually, to keep things simple. Output looks like:

53

/// manually, to keep things simple. Output looks like:

125

///

54

///

126

///xml```

55

///xml```

127

/// <environment_context>

56

/// <environment_context>

128

/// <cwd>...</cwd>

57

/// <cwd>...</cwd>

129

/// <approval_policy>...</approval_policy>

130

/// <sandbox_mode>...</sandbox_mode>

131

/// <writable_roots>...</writable_roots>

132

/// <network_access>...</network_access>

133

/// <shell>...</shell>

58

/// <shell>...</shell>

134

/// </environment_context>

59

/// </environment_context>

135

``` /// ``````

60

``` /// ``````

136

pub fn serialize_to_xml(self) -> String {

61

pub fn serialize_to_xml(self) -> String {

137

let mut lines = vec![ENVIRONMENT_CONTEXT_OPEN_TAG.to_string()];

62

let mut lines = vec![ENVIRONMENT_CONTEXT_OPEN_TAG.to_string()];

138

if let Some(cwd) = self.cwd {

63

if let Some(cwd) = self.cwd {

139

lines.push(format!(" <cwd>{}</cwd>", cwd.to_string_lossy()));

64

lines.push(format!(" <cwd>{}</cwd>", cwd.to_string_lossy()));

140

}

65

}

141

if let Some(approval_policy) = self.approval_policy {

142

lines.push(format!(

143

" <approval_policy>{approval_policy}</approval_policy>"

144

));

145

}

146

if let Some(sandbox_mode) = self.sandbox_mode {

147

lines.push(format!(" <sandbox_mode>{sandbox_mode}</sandbox_mode>"));

148

}

149

if let Some(network_access) = self.network_access {

150

lines.push(format!(

151

" <network_access>{network_access}</network_access>"

152

));

153

}

154

if let Some(writable_roots) = self.writable_roots {

155

lines.push(" <writable_roots>".to_string());

156

for writable_root in writable_roots {

157

lines.push(format!(

158

" <root>{}</root>",

159

writable_root.to_string_lossy()

160

));

161

}

162

lines.push(" </writable_roots>".to_string());

163

}

164

66

165

let shell_name = self.shell.name();

67

let shell_name = self.shell.name();

166

lines.push(format!(" <shell>{shell_name}</shell>"));

68

lines.push(format!(" <shell>{shell_name}</shell>"));

167

lines.push(ENVIRONMENT_CONTEXT_CLOSE_TAG.to_string());

69

lines.push(ENVIRONMENT_CONTEXT_CLOSE_TAG.to_string());

168

lines.join("\n")

70

lines.join("\n")

169

}

71

}

170

}

72

}

171

73

172

impl From<EnvironmentContext> for ResponseItem {

74

impl From<EnvironmentContext> for ResponseItem {

All 9 lines

All 9 lines

182

}

84

}

183

85

184

#[cfg(test)]

86

#[cfg(test)]

185

mod tests {

87

mod tests {

186

use crate::shell::ShellType;

88

use crate::shell::ShellType;

187

89

188

use super::*;

90

use super::*;

189

use core_test_support::test_path_buf;

91

use core_test_support::test_path_buf;

190

use core_test_support::test_tmp_path_buf;

191

use pretty_assertions::assert_eq;

92

use pretty_assertions::assert_eq;

192

93

193

fn fake_shell() -> Shell {

94

fn fake_shell() -> Shell {

All 5 lines

All 5 lines

199

}

100

}

200

101

201

fn workspace_write_policy(writable_roots: Vec<&str>, network_access: bool) -> SandboxPolicy {

202

SandboxPolicy::WorkspaceWrite {

203

writable_roots: writable_roots

204

.into_iter()

205

.map(|s| AbsolutePathBuf::try_from(s).unwrap())

206

.collect(),

207

network_access,

208

exclude_tmpdir_env_var: false,

209

exclude_slash_tmp: false,

210

}

211

}

212

213

#[test]

102

#[test]

214

fn serialize_workspace_write_environment_context() {

103

fn serialize_workspace_write_environment_context() {

215

let cwd = test_path_buf("/repo");

104

let cwd = test_path_buf("/repo");

216

let writable_root = test_tmp_path_buf();

105

let context = EnvironmentContext::new(Some(cwd.clone()), fake_shell());

217

let cwd_str = cwd.to_str().expect("cwd is valid utf-8");

218

let writable_root_str = writable_root

219

.to_str()

220

.expect("writable root is valid utf-8");

221

let context = EnvironmentContext::new(

222

Some(cwd.clone()),

223

Some(AskForApproval::OnRequest),

224

Some(workspace_write_policy(

225

vec![cwd_str, writable_root_str],

226

false,

227

)),

228

fake_shell(),

229

);

230

106

231

let expected = format!(

107

let expected = format!(

232

r#"<environment_context>

108

r#"<environment_context>

233

<cwd>{cwd}</cwd>

109

<cwd>{cwd}</cwd>

234

<approval_policy>on-request</approval_policy>

235

<sandbox_mode>workspace-write</sandbox_mode>

236

<network_access>restricted</network_access>

237

<writable_roots>

238

<root>{cwd}</root>

239

<root>{writable_root}</root>

240

</writable_roots>

241

<shell>bash</shell>

110

<shell>bash</shell>

242

</environment_context>"#,

111

</environment_context>"#,

243

cwd = cwd.display(),

112

cwd = cwd.display(),

244

writable_root = writable_root.display(),

245

);

113

);

246

114

247

assert_eq!(context.serialize_to_xml(), expected);

115

assert_eq!(context.serialize_to_xml(), expected);

248

}

116

}

249

117

250

#[test]

118

#[test]

251

fn serialize_read_only_environment_context() {

119

fn serialize_read_only_environment_context() {

252

let context = EnvironmentContext::new(

120

let context = EnvironmentContext::new(None, fake_shell());

253

None,

254

Some(AskForApproval::Never),

255

Some(SandboxPolicy::ReadOnly),

256

fake_shell(),

257

);

258

121

259

let expected = r#"<environment_context>

122

let expected = r#"<environment_context>

260

<approval_policy>never</approval_policy>

261

<sandbox_mode>read-only</sandbox_mode>

262

<network_access>restricted</network_access>

263

<shell>bash</shell>

123

<shell>bash</shell>

264

</environment_context>"#;

124

</environment_context>"#;

265

125

266

assert_eq!(context.serialize_to_xml(), expected);

126

assert_eq!(context.serialize_to_xml(), expected);

267

}

127

}

268

128

269

#[test]

129

#[test]

270

fn serialize_external_sandbox_environment_context() {

130

fn serialize_external_sandbox_environment_context() {

271

let context = EnvironmentContext::new(

131

let context = EnvironmentContext::new(None, fake_shell());

272

None,

273

Some(AskForApproval::OnRequest),

274

Some(SandboxPolicy::ExternalSandbox {

275

network_access: NetworkAccess::Enabled,

276

}),

277

fake_shell(),

278

);

279

132

280

let expected = r#"<environment_context>

133

let expected = r#"<environment_context>

281

<approval_policy>on-request</approval_policy>

282

<sandbox_mode>danger-full-access</sandbox_mode>

283

<network_access>enabled</network_access>

284

<shell>bash</shell>

134

<shell>bash</shell>

285

</environment_context>"#;

135

</environment_context>"#;

286

136

287

assert_eq!(context.serialize_to_xml(), expected);

137

assert_eq!(context.serialize_to_xml(), expected);

288

}

138

}

289

139

290

#[test]

140

#[test]

291

fn serialize_external_sandbox_with_restricted_network_environment_context() {

141

fn serialize_external_sandbox_with_restricted_network_environment_context() {

292

let context = EnvironmentContext::new(

142

let context = EnvironmentContext::new(None, fake_shell());

293

None,

294

Some(AskForApproval::OnRequest),

295

Some(SandboxPolicy::ExternalSandbox {

296

network_access: NetworkAccess::Restricted,

297

}),

298

fake_shell(),

299

);

300

143

301

let expected = r#"<environment_context>

144

let expected = r#"<environment_context>

302

<approval_policy>on-request</approval_policy>

303

<sandbox_mode>danger-full-access</sandbox_mode>

304

<network_access>restricted</network_access>

305

<shell>bash</shell>

145

<shell>bash</shell>

306

</environment_context>"#;

146

</environment_context>"#;

307

147

308

assert_eq!(context.serialize_to_xml(), expected);

148

assert_eq!(context.serialize_to_xml(), expected);

309

}

149

}

310

150

311

#[test]

151

#[test]

312

fn serialize_full_access_environment_context() {

152

fn serialize_full_access_environment_context() {

313

let context = EnvironmentContext::new(

153

let context = EnvironmentContext::new(None, fake_shell());

314

None,

315

Some(AskForApproval::OnFailure),

316

Some(SandboxPolicy::DangerFullAccess),

317

fake_shell(),

318

);

319

154

320

let expected = r#"<environment_context>

155

let expected = r#"<environment_context>

321

<approval_policy>on-failure</approval_policy>

322

<sandbox_mode>danger-full-access</sandbox_mode>

323

<network_access>enabled</network_access>

324

<shell>bash</shell>

156

<shell>bash</shell>

325

</environment_context>"#;

157

</environment_context>"#;

326

158

327

assert_eq!(context.serialize_to_xml(), expected);

159

assert_eq!(context.serialize_to_xml(), expected);

328

}

160

}

329

161

330

#[test]

162

#[test]

331

fn equals_except_shell_compares_approval_policy() {

163

fn equals_except_shell_compares_cwd() {

332

// Approval policy

164

let context1 = EnvironmentContext::new(Some(PathBuf::from("/repo")), fake_shell());

333

let context1 = EnvironmentContext::new(

165

let context2 = EnvironmentContext::new(Some(PathBuf::from("/repo")), fake_shell());

334

Some(PathBuf::from("/repo")),

166

assert!(context1.equals_except_shell(&context2));

335

Some(AskForApproval::OnRequest),

336

Some(workspace_write_policy(vec!["/repo"], false)),

337

fake_shell(),

338

);

339

let context2 = EnvironmentContext::new(

340

Some(PathBuf::from("/repo")),

341

Some(AskForApproval::Never),

342

Some(workspace_write_policy(vec!["/repo"], true)),

343

fake_shell(),

344

);

345

assert!(!context1.equals_except_shell(&context2));

346

}

167

}

347

168

348

#[test]

169

#[test]

349

fn equals_except_shell_compares_sandbox_policy() {

170

fn equals_except_shell_ignores_sandbox_policy() {

350

let context1 = EnvironmentContext::new(

171

let context1 = EnvironmentContext::new(Some(PathBuf::from("/repo")), fake_shell());

351

Some(PathBuf::from("/repo")),

172

let context2 = EnvironmentContext::new(Some(PathBuf::from("/repo")), fake_shell());

352

Some(AskForApproval::OnRequest),

353

Some(SandboxPolicy::new_read_only_policy()),

354

fake_shell(),

355

);

356

let context2 = EnvironmentContext::new(

357

Some(PathBuf::from("/repo")),

358

Some(AskForApproval::OnRequest),

359

Some(SandboxPolicy::new_workspace_write_policy()),

360

fake_shell(),

361

);

362

173

363

assert!(!context1.equals_except_shell(&context2));

174

assert!(context1.equals_except_shell(&context2));

364

}

175

}

365

176

366

#[test]

177

#[test]

367

fn equals_except_shell_compares_workspace_write_policy() {

178

fn equals_except_shell_compares_cwd_differences() {

368

let context1 = EnvironmentContext::new(

179

let context1 = EnvironmentContext::new(Some(PathBuf::from("/repo1")), fake_shell());

369

Some(PathBuf::from("/repo")),

180

let context2 = EnvironmentContext::new(Some(PathBuf::from("/repo2")), fake_shell());

370

Some(AskForApproval::OnRequest),

371

Some(workspace_write_policy(vec!["/repo", "/tmp", "/var"], false)),

372

fake_shell(),

373

);

374

let context2 = EnvironmentContext::new(

375

Some(PathBuf::from("/repo")),

376

Some(AskForApproval::OnRequest),

377

Some(workspace_write_policy(vec!["/repo", "/tmp"], true)),

378

fake_shell(),

379

);

380

181

381

assert!(!context1.equals_except_shell(&context2));

182

assert!(!context1.equals_except_shell(&context2));

382

}

183

}

383

184

384

#[test]

185

#[test]

385

fn equals_except_shell_ignores_shell() {

186

fn equals_except_shell_ignores_shell() {

386

let context1 = EnvironmentContext::new(

187

let context1 = EnvironmentContext::new(

387

Some(PathBuf::from("/repo")),

188

Some(PathBuf::from("/repo")),

388

Some(AskForApproval::OnRequest),

389

Some(workspace_write_policy(vec!["/repo"], false)),

390

Shell {

189

Shell {

391

shell_type: ShellType::Bash,

190

shell_type: ShellType::Bash,

392

shell_path: "/bin/bash".into(),

191

shell_path: "/bin/bash".into(),

393

shell_snapshot: None,

192

shell_snapshot: None,

394

},

193

},

395

);

194

);

396

let context2 = EnvironmentContext::new(

195

let context2 = EnvironmentContext::new(

397

Some(PathBuf::from("/repo")),

196

Some(PathBuf::from("/repo")),

398

Some(AskForApproval::OnRequest),

399

Some(workspace_write_policy(vec!["/repo"], false)),

400

Shell {

197

Shell {

401

shell_type: ShellType::Zsh,

198

shell_type: ShellType::Zsh,

402

shell_path: "/bin/zsh".into(),

199

shell_path: "/bin/zsh".into(),

403

shell_snapshot: None,

200

shell_snapshot: None,

404

},

201

},

405

);

202

);

406

203

407

assert!(context1.equals_except_shell(&context2));

204

assert!(context1.equals_except_shell(&context2));

408

}

205

}

409

}

206

}

4Removed Static Prompt Sections

0 / 7

gpt-5.1-codex-max_prompt.mdcodex-rs/core

−37

Mark as viewed

5 linesAll 20 lines5 lines

5 linesAll 20 lines5 lines

21

## Plan tool

21

## Plan tool

All 6 lines

All 6 lines

28

## Codex CLI harness, sandboxing, and approvals

29

30

The Codex CLI harness supports several different configurations for sandboxing and escalation approvals that the user can choose from.

31

32

Filesystem sandboxing defines which files can be read or written. The options for `sandbox_mode` are:

33

- **read-only**: The sandbox only permits reading files.

34

- **workspace-write**: The sandbox permits reading files, and editing files in `cwd` and `writable_roots`. Editing files in other directories requires approval.

35

- **danger-full-access**: No filesystem sandboxing - all commands are permitted.

36

37

Network sandboxing defines whether network can be accessed without approval. Options for `network_access` are:

38

- **restricted**: Requires approval

39

- **enabled**: No approval needed

40

41

Approvals are your mechanism to get user consent to run shell commands without the sandbox. Possible configuration options for `approval_policy` are

42

- **untrusted**: The harness will escalate most commands for user approval, apart from a limited allowlist of safe "read" commands.

43

- **on-failure**: The harness will allow all commands to run in the sandbox (if enabled), and failures will be escalated to the user for approval to run again without the sandbox.

44

- **on-request**: Commands will be run in the sandbox by default, and you can specify in your tool call if you want to escalate a command to run without sandboxing. (Note that this mode is not always available. If it is, you'll see parameters for it in the `shell` command description.)

45

- **never**: This is a non-interactive mode where you may NEVER ask the user for approval to run commands. Instead, you must always persist and work around constraints to solve the task for the user. You MUST do your utmost best to finish the task and validate your work before yielding. If this mode is paired with `danger-full-access`, take advantage of it to deliver the best outcome for the user. Further, in this mode, your default testing philosophy is overridden: Even if you don't see local patterns for testing, you may add tests and scripts to validate your work. Just remove them before yielding.

46

47

When you are running with `approval_policy == on-request`, and sandboxing enabled, here are scenarios where you'll need to request approval:

48

- You need to run a command that writes to a directory that requires it (e.g. running tests that write to /var)

49

- You need to run a GUI app (e.g., open/xdg-open/osascript) to open browsers or files.

50

- You are running sandboxed and need to run a command that requires network access (e.g. installing packages)

51

52

- You are about to take a potentially destructive action such as an `rm` or `git reset` that the user did not explicitly ask for

53

- (for all of these, you should weigh alternative paths that do not require approval)

54

55

When `sandbox_mode` is set to read-only, you'll need to request approval for any command that isn't a read.

56

57

You will be told what filesystem sandboxing, network sandboxing, and approval mode are active in a developer or user message. If you are not told about this, assume that you are running with workspace-write, network sandboxing enabled, and approval on-failure.

58

59

Although they introduce friction to the user because your work is paused until the user responds, you should leverage them when necessary to accomplish important work. If the completing the task requires escalated permissions, Do not let these settings or the sandbox deter you from attempting to accomplish the user's task unless it is set to "never", in which case never ask for approvals.

60

61

When requesting approval to execute a command that will require escalated privileges:

62

`` - Provide the sandbox_permissions parameter with the value `"require_escalated"```

63

- Include a short, 1 sentence explanation for why you need escalated permissions in the justification parameter

64

65

## Special user requests

28

## Special user requests

5 linesAll 52 lines5 lines

5 linesAll 52 lines5 lines

gpt-5.2-codex_prompt.mdcodex-rs/core

−37

Mark as viewed

1

You are Codex, based on GPT-5. You are running as a coding agent in the Codex CLI on a user's computer.

1

You are Codex, based on GPT-5. You are running as a coding agent in the Codex CLI on a user's computer.

2

2

3

## General

3

## General

4

4

5

- When searching for text or files, prefer using `rg` or `rg --files` respectively because `rg` is much faster than alternatives like `grep`. (If the `rg` command is not found, then use alternatives.)

5

6

6

7

## Editing constraints

7

## Editing constraints

5 linesAll 13 lines5 lines

5 linesAll 13 lines5 lines

21

## Plan tool

21

## Plan tool

All 6 lines

All 6 lines

28

## Codex CLI harness, sandboxing, and approvals

29

30

The Codex CLI harness supports several different configurations for sandboxing and escalation approvals that the user can choose from.

31

32

Filesystem sandboxing defines which files can be read or written. The options for `sandbox_mode` are:

33

- **read-only**: The sandbox only permits reading files.

34

35

- **danger-full-access**: No filesystem sandboxing - all commands are permitted.

36

37

Network sandboxing defines whether network can be accessed without approval. Options for `network_access` are:

38

- **restricted**: Requires approval

39

- **enabled**: No approval needed

40

41

Approvals are your mechanism to get user consent to run shell commands without the sandbox. Possible configuration options for `approval_policy` are

42

- **untrusted**: The harness will escalate most commands for user approval, apart from a limited allowlist of safe "read" commands.

43

44

45

46

47

When you are running with `approval_policy == on-request`, and sandboxing enabled, here are scenarios where you'll need to request approval:

48

- You need to run a command that writes to a directory that requires it (e.g. running tests that write to /var)

49

- You need to run a GUI app (e.g., open/xdg-open/osascript) to open browsers or files.

50

- You are running sandboxed and need to run a command that requires network access (e.g. installing packages)

51

52

- You are about to take a potentially destructive action such as an `rm` or `git reset` that the user did not explicitly ask for

53

- (for all of these, you should weigh alternative paths that do not require approval)

54

55

When `sandbox_mode` is set to read-only, you'll need to request approval for any command that isn't a read.

56

57

58

59

60

61

When requesting approval to execute a command that will require escalated privileges:

62

`` - Provide the sandbox_permissions parameter with the value `"require_escalated"```

63

- Include a short, 1 sentence explanation for why you need escalated permissions in the justification parameter

64

65

## Special user requests

28

## Special user requests

5 linesAll 52 lines5 lines

5 linesAll 52 lines5 lines

gpt_5_1_prompt.mdcodex-rs/core

−37

Mark as viewed

5 linesAll 135 lines5 lines

5 linesAll 135 lines5 lines

136

## Task execution

136

## Task execution

5 linesAll 24 lines5 lines

5 linesAll 24 lines5 lines

161

161

162

## Codex CLI harness, sandboxing, and approvals

163

164

The Codex CLI harness supports several different configurations for sandboxing and escalation approvals that the user can choose from.

165

166

Filesystem sandboxing defines which files can be read or written. The options for `sandbox_mode` are:

1

Move167–169

167

- **read-only**: The sandbox only permits reading files.

168

169

- **danger-full-access**: No filesystem sandboxing - all commands are permitted.

170

171

Network sandboxing defines whether network can be accessed without approval. Options for `network_access` are:

172

- **restricted**: Requires approval

173

- **enabled**: No approval needed

174

1

Move175–179

175

Approvals are your mechanism to get user consent to run shell commands without the sandbox. Possible configuration options for `approval_policy` are

176

- **untrusted**: The harness will escalate most commands for user approval, apart from a limited allowlist of safe "read" commands.

177

178

- **on-request**: Commands will be run in the sandbox by default, and you can specify in your tool call if you want to escalate a command to run without sandboxing. (Note that this mode is not always available. If it is, you'll see parameters for escalating in the tool definition.)

179

180

181

When you are running with `approval_policy == on-request`, and sandboxing enabled, here are scenarios where you'll need to request approval:

182

- You need to run a command that writes to a directory that requires it (e.g. running tests that write to /var)

183

- You need to run a GUI app (e.g., open/xdg-open/osascript) to open browsers or files.

184

- You are running sandboxed and need to run a command that requires network access (e.g. installing packages)

185

- If you run a command that is important to solving the user's query, but it fails because of sandboxing, rerun the command with approval. ALWAYS proceed to use the `sandbox_permissions` and `justification` parameters. Within this harness, prefer requesting approval via the tool over asking in natural language.

186

- You are about to take a potentially destructive action such as an `rm` or `git reset` that the user did not explicitly ask for

187

- (for all of these, you should weigh alternative paths that do not require approval)

188

189

When `sandbox_mode` is set to read-only, you'll need to request approval for any command that isn't a read.

190

191

192

193

194

195

When requesting approval to execute a command that will require escalated privileges:

196

`` - Provide the sandbox_permissions parameter with the value `"require_escalated"```

197

- Include a short, 1 sentence explanation for why you need escalated permissions in the justification parameter

198

199

## Validating your work

162

## Validating your work

5 linesAll 169 lines5 lines

5 linesAll 169 lines5 lines

gpt_5_2_prompt.mdcodex-rs/core

−37

Mark as viewed

5 linesAll 108 lines5 lines

5 linesAll 108 lines5 lines

109

## Task execution

109

## Task execution

5 linesAll 24 lines5 lines

5 linesAll 24 lines5 lines

134

- NEVER output inline citations like "【F:README.md†L5-L14】" in your outputs. The CLI is not able to render these so they will just be broken in the UI. Instead, if you output valid filepaths, users will be able to click on them to open the files in their editor.

134

135

135

136

## Codex CLI harness, sandboxing, and approvals

137

138

The Codex CLI harness supports several different configurations for sandboxing and escalation approvals that the user can choose from.

139

140

Filesystem sandboxing defines which files can be read or written. The options for `sandbox_mode` are:

141

- **read-only**: The sandbox only permits reading files.

142

143

- **danger-full-access**: No filesystem sandboxing - all commands are permitted.

144

145

Network sandboxing defines whether network can be accessed without approval. Options for `network_access` are:

146

- **restricted**: Requires approval

147

- **enabled**: No approval needed

148

149

Approvals are your mechanism to get user consent to run shell commands without the sandbox. Possible configuration options for `approval_policy` are

150

- **untrusted**: The harness will escalate most commands for user approval, apart from a limited allowlist of safe "read" commands.

151

152

153

154

155

When you are running with `approval_policy == on-request`, and sandboxing enabled, here are scenarios where you'll need to request approval:

156

- You need to run a command that writes to a directory that requires it (e.g. running tests that write to /var)

157

- You need to run a GUI app (e.g., open/xdg-open/osascript) to open browsers or files.

158

- You are running sandboxed and need to run a command that requires network access (e.g. installing packages)

159

160

- You are about to take a potentially destructive action such as an `rm` or `git reset` that the user did not explicitly ask for

161

- (for all of these, you should weigh alternative paths that do not require approval)

162

163

When `sandbox_mode` is set to read-only, you'll need to request approval for any command that isn't a read.

164

165

166

167

168

169

When requesting approval to execute a command that will require escalated privileges:

170

`` - Provide the sandbox_permissions parameter with the value `"require_escalated"```

171

- Include a short, 1 sentence explanation for why you need escalated permissions in the justification parameter

172

173

## Validating your work

136

## Validating your work

5 linesAll 162 lines5 lines

5 linesAll 162 lines5 lines

gpt_5_codex_prompt.mdcodex-rs/core

−37

Mark as viewed

5 linesAll 20 lines5 lines

5 linesAll 20 lines5 lines

21

## Plan tool

21

## Plan tool

All 6 lines

All 6 lines

28

## Codex CLI harness, sandboxing, and approvals

29

30

The Codex CLI harness supports several different configurations for sandboxing and escalation approvals that the user can choose from.

31

32

Filesystem sandboxing defines which files can be read or written. The options for `sandbox_mode` are:

33

- **read-only**: The sandbox only permits reading files.

34

35

- **danger-full-access**: No filesystem sandboxing - all commands are permitted.

36

37

Network sandboxing defines whether network can be accessed without approval. Options for `network_access` are:

38

- **restricted**: Requires approval

39

- **enabled**: No approval needed

40

41

Approvals are your mechanism to get user consent to run shell commands without the sandbox. Possible configuration options for `approval_policy` are

42

- **untrusted**: The harness will escalate most commands for user approval, apart from a limited allowlist of safe "read" commands.

43

44

45

46

47

When you are running with `approval_policy == on-request`, and sandboxing enabled, here are scenarios where you'll need to request approval:

48

- You need to run a command that writes to a directory that requires it (e.g. running tests that write to /var)

49

- You need to run a GUI app (e.g., open/xdg-open/osascript) to open browsers or files.

50

- You are running sandboxed and need to run a command that requires network access (e.g. installing packages)

51

52

- You are about to take a potentially destructive action such as an `rm` or `git reset` that the user did not explicitly ask for

53

- (for all of these, you should weigh alternative paths that do not require approval)

54

55

When `sandbox_mode` is set to read-only, you'll need to request approval for any command that isn't a read.

56

57

58

59

60

61

When requesting approval to execute a command that will require escalated privileges:

62

`` - Provide the sandbox_permissions parameter with the value `"require_escalated"```

63

- Include a short, 1 sentence explanation for why you need escalated permissions in the justification parameter

64

65

## Special user requests

28

## Special user requests

5 linesAll 40 lines5 lines

5 linesAll 40 lines5 lines

prompt.mdcodex-rs/core

−35

Mark as viewed

Informational

5 linesAll 122 lines5 lines

5 linesAll 122 lines5 lines

123

## Task execution

123

## Task execution

5 linesAll 23 lines5 lines

5 linesAll 23 lines5 lines

147

147

148

148

149

## Sandbox and approvals

150

151

The Codex CLI harness supports several different sandboxing, and approval configurations that the user can choose from.

152

153

Filesystem sandboxing prevents you from editing files without user approval. The options are:

154

155

- **read-only**: You can only read files.

156

- **workspace-write**: You can read files. You can write to files in your workspace folder, but not outside it.

157

- **danger-full-access**: No filesystem sandboxing.

158

159

Network sandboxing prevents you from accessing network without approval. Options are

160

161

- **restricted**

162

- **enabled**

163

164

Approvals are your mechanism to get user consent to perform more privileged actions. Although they introduce friction to the user because your work is paused until the user responds, you should leverage them to accomplish your important work. Do not let these settings or the sandbox deter you from attempting to accomplish the user's task. Approval options are

165

166

- **untrusted**: The harness will escalate most commands for user approval, apart from a limited allowlist of safe "read" commands.

167

168

169

- **never**: This is a non-interactive mode where you may NEVER ask the user for approval to run commands. Instead, you must always persist and work around constraints to solve the task for the user. You MUST do your utmost best to finish the task and validate your work before yielding. If this mode is pared with `danger-full-access`, take advantage of it to deliver the best outcome for the user. Further, in this mode, your default testing philosophy is overridden: Even if you don't see local patterns for testing, you may add tests and scripts to validate your work. Just remove them before yielding.

170

171

When you are running with approvals `on-request`, and sandboxing enabled, here are scenarios where you'll need to request approval:

172

173

- You need to run a command that writes to a directory that requires it (e.g. running tests that write to /tmp)

174

- You need to run a GUI app (e.g., open/xdg-open/osascript) to open browsers or files.

175

- You are running sandboxed and need to run a command that requires network access (e.g. installing packages)

176

- If you run a command that is important to solving the user's query, but it fails because of sandboxing, rerun the command with approval.

177

- You are about to take a potentially destructive action such as an `rm` or `git reset` that the user did not explicitly ask for

178

- (For all of these, you should weigh alternative paths that do not require approval.)

179

180

Note that when sandboxing is set to read-only, you'll need to request approval for any command that isn't a read.

181

182

You will be told what filesystem sandboxing, network sandboxing, and approval mode are active in a developer or user message. If you are not told about this, assume that you are running with workspace-write, network sandboxing ON, and approval on-failure.

183

184

## Validating your work

149

## Validating your work

5 linesAll 126 lines5 lines

5 linesAll 126 lines5 lines

prompt_with_apply_patch_instructions.mdcodex-rs/core

−35

Mark as viewed

5 linesAll 122 lines5 lines

5 linesAll 122 lines5 lines

123

## Task execution

123

## Task execution

5 linesAll 25 lines5 lines

5 linesAll 25 lines5 lines

149

## Sandbox and approvals

150

151

The Codex CLI harness supports several different sandboxing, and approval configurations that the user can choose from.

152

153

Filesystem sandboxing prevents you from editing files without user approval. The options are:

154

155

- **read-only**: You can only read files.

156

- **workspace-write**: You can read files. You can write to files in your workspace folder, but not outside it.

157

- **danger-full-access**: No filesystem sandboxing.

158

159

Network sandboxing prevents you from accessing network without approval. Options are

160

161

- **restricted**

162

- **enabled**

163

164

165

166

- **untrusted**: The harness will escalate most commands for user approval, apart from a limited allowlist of safe "read" commands.

167

168

169

170

171

When you are running with approvals `on-request`, and sandboxing enabled, here are scenarios where you'll need to request approval:

172

173

- You need to run a command that writes to a directory that requires it (e.g. running tests that write to /tmp)

174

- You need to run a GUI app (e.g., open/xdg-open/osascript) to open browsers or files.

175

- You are running sandboxed and need to run a command that requires network access (e.g. installing packages)

176

- If you run a command that is important to solving the user's query, but it fails because of sandboxing, rerun the command with approval.

177

- You are about to take a potentially destructive action such as an `rm` or `git reset` that the user did not explicitly ask for

178

- (For all of these, you should weigh alternative paths that do not require approval.)

179

180

Note that when sandboxing is set to read-only, you'll need to request approval for any command that isn't a read.

181

182

183

184

## Validating your work

149

## Validating your work

5 linesAll 202 lines5 lines

5 linesAll 202 lines5 lines

5New permissions_messages Test Suite

0 / 2

mod.rscodex-rs/core/tests/suite

+1

Mark as viewed

5 linesAll 15 lines5 lines

5 linesAll 15 lines5 lines

16

#[cfg(not(target_os = "windows"))]

16

#[cfg(not(target_os = "windows"))]

17

mod abort_tasks;

17

mod abort_tasks;

5 linesAll 24 lines5 lines

5 linesAll 24 lines5 lines

42

mod otel;

42

mod otel;

43

mod permissions_messages;

43

mod prompt_caching;

44

mod prompt_caching;

5 linesAll 27 lines5 lines

5 linesAll 27 lines5 lines

permissions_messages.rscodex-rs/core/tests/suite

+448

AddedMark as viewed

1

use anyhow::Result;

2

use codex_core::config::Constrained;

3

use codex_core::protocol::AskForApproval;

4

use codex_core::protocol::EventMsg;

5

use codex_core::protocol::Op;

6

use codex_core::protocol::SandboxPolicy;

7

use codex_protocol::user_input::UserInput;

8

use codex_utils_absolute_path::AbsolutePathBuf;

9

use core_test_support::responses::ev_completed;

10

use core_test_support::responses::ev_response_created;

11

use core_test_support::responses::mount_sse_once;

12

use core_test_support::responses::sse;

13

use core_test_support::responses::start_mock_server;

14

use core_test_support::skip_if_no_network;

15

use core_test_support::test_codex::test_codex;

16

use core_test_support::wait_for_event;

17

use pretty_assertions::assert_eq;

18

use std::collections::HashSet;

19

use tempfile::TempDir;

20

21

fn permissions_texts(input: &[serde_json::Value]) -> Vec<String> {

22

input

23

.iter()

24

.filter_map(|item| {

25

let role = item.get("role")?.as_str()?;

26

if role != "developer" {

27

return None;

28

}

29

let text = item

30

.get("content")?

31

.as_array()?

32

.first()?

33

.get("text")?

34

.as_str()?;

35

if text.contains("`approval_policy`") {

36

Some(text.to_string())

37

} else {

38

None

39

}

40

})

41

.collect()

42

}

43

44

fn sse_completed(id: &str) -> String {

45

sse(vec![ev_response_created(id), ev_completed(id)])

46

}

47

48

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

49

async fn permissions_message_sent_once_on_start() -> Result<()> {

50

skip_if_no_network!(Ok(()));

51

52

let server = start_mock_server().await;

53

let req = mount_sse_once(&server, sse_completed("resp-1")).await;

54

55

let mut builder = test_codex().with_config(move |config| {

56

config.approval_policy = Constrained::allow_any(AskForApproval::OnRequest);

57

});

58

let test = builder.build(&server).await?;

59

60

test.codex

61

.submit(Op::UserInput {

62

items: vec![UserInput::Text {

63

text: "hello".into(),

64

}],

65

final_output_json_schema: None,

66

})

67

.await?;

68

wait_for_event(&test.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;

69

70

let request = req.single_request();

71

let body = request.body_json();

72

let input = body["input"].as_array().expect("input array");

73

let permissions = permissions_texts(input);

74

assert_eq!(permissions.len(), 1);

75

76

Ok(())

77

}

78

79

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

80

async fn permissions_message_added_on_override_change() -> Result<()> {

81

skip_if_no_network!(Ok(()));

82

83

let server = start_mock_server().await;

84

let req1 = mount_sse_once(&server, sse_completed("resp-1")).await;

85

let req2 = mount_sse_once(&server, sse_completed("resp-2")).await;

86

87

let mut builder = test_codex().with_config(move |config| {

88

config.approval_policy = Constrained::allow_any(AskForApproval::OnRequest);

89

});

90

let test = builder.build(&server).await?;

91

92

test.codex

93

.submit(Op::UserInput {

94

items: vec![UserInput::Text {

95

text: "hello 1".into(),

96

}],

97

final_output_json_schema: None,

98

})

99

.await?;

100

wait_for_event(&test.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;

101

102

test.codex

103

.submit(Op::OverrideTurnContext {

104

cwd: None,

105

approval_policy: Some(AskForApproval::Never),

106

sandbox_policy: None,

107

model: None,

108

effort: None,

109

summary: None,

110

})

111

.await?;

112

113

test.codex

114

.submit(Op::UserInput {

115

items: vec![UserInput::Text {

116

text: "hello 2".into(),

117

}],

118

final_output_json_schema: None,

119

})

120

.await?;

121

wait_for_event(&test.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;

122

123

let body1 = req1.single_request().body_json();

124

let body2 = req2.single_request().body_json();

125

let input1 = body1["input"].as_array().expect("input array");

126

let input2 = body2["input"].as_array().expect("input array");

127

let permissions_1 = permissions_texts(input1);

128

let permissions_2 = permissions_texts(input2);

129

130

assert_eq!(permissions_1.len(), 1);

131

assert_eq!(permissions_2.len(), 2);

132

let unique = permissions_2.into_iter().collect::<HashSet<String>>();

133

assert_eq!(unique.len(), 2);

134

135

Ok(())

136

}

137

138

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

139

async fn permissions_message_not_added_when_no_change() -> Result<()> {

140

skip_if_no_network!(Ok(()));

141

142

let server = start_mock_server().await;

143

let req1 = mount_sse_once(&server, sse_completed("resp-1")).await;

144

let req2 = mount_sse_once(&server, sse_completed("resp-2")).await;

145

146

let mut builder = test_codex().with_config(move |config| {

147

config.approval_policy = Constrained::allow_any(AskForApproval::OnRequest);

148

});

149

let test = builder.build(&server).await?;

150

151

test.codex

152

.submit(Op::UserInput {

153

items: vec![UserInput::Text {

154

text: "hello 1".into(),

155

}],

156

final_output_json_schema: None,

157

})

158

.await?;

159

wait_for_event(&test.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;

160

161

test.codex

162

.submit(Op::UserInput {

163

items: vec![UserInput::Text {

164

text: "hello 2".into(),

165

}],

166

final_output_json_schema: None,

167

})

168

.await?;

169

wait_for_event(&test.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;

170

171

let body1 = req1.single_request().body_json();

172

let body2 = req2.single_request().body_json();

173

let input1 = body1["input"].as_array().expect("input array");

174

let input2 = body2["input"].as_array().expect("input array");

175

let permissions_1 = permissions_texts(input1);

176

let permissions_2 = permissions_texts(input2);

177

178

assert_eq!(permissions_1.len(), 1);

179

assert_eq!(permissions_2.len(), 1);

180

assert_eq!(permissions_1, permissions_2);

181

182

Ok(())

183

}

184

185

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

186

async fn resume_replays_permissions_messages() -> Result<()> {

187

skip_if_no_network!(Ok(()));

188

189

let server = start_mock_server().await;

190

let _req1 = mount_sse_once(&server, sse_completed("resp-1")).await;

191

let _req2 = mount_sse_once(&server, sse_completed("resp-2")).await;

192

let req3 = mount_sse_once(&server, sse_completed("resp-3")).await;

193

194

let mut builder = test_codex().with_config(|config| {

195

config.approval_policy = Constrained::allow_any(AskForApproval::OnRequest);

196

});

197

let initial = builder.build(&server).await?;

198

let rollout_path = initial.session_configured.rollout_path.clone();

199

let home = initial.home.clone();

200

201

initial

202

.codex

203

.submit(Op::UserInput {

204

items: vec![UserInput::Text {

205

text: "hello 1".into(),

206

}],

207

final_output_json_schema: None,

208

})

209

.await?;

210

wait_for_event(&initial.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;

211

212

initial

213

.codex

214

.submit(Op::OverrideTurnContext {

215

cwd: None,

216

approval_policy: Some(AskForApproval::Never),

217

sandbox_policy: None,

218

model: None,

219

effort: None,

220

summary: None,

221

})

222

.await?;

223

224

initial

225

.codex

226

.submit(Op::UserInput {

227

items: vec![UserInput::Text {

228

text: "hello 2".into(),

229

}],

230

final_output_json_schema: None,

231

})

232

.await?;

233

wait_for_event(&initial.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;

234

235

let resumed = builder.resume(&server, home, rollout_path).await?;

236

resumed

237

.codex

238

.submit(Op::UserInput {

239

items: vec![UserInput::Text {

240

text: "after resume".into(),

241

}],

242

final_output_json_schema: None,

243

})

244

.await?;

245

wait_for_event(&resumed.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;

246

247

let body3 = req3.single_request().body_json();

248

let input = body3["input"].as_array().expect("input array");

249

let permissions = permissions_texts(input);

250

assert_eq!(permissions.len(), 3);

251

let unique = permissions.into_iter().collect::<HashSet<String>>();

252

assert_eq!(unique.len(), 2);

253

254

Ok(())

255

}

256

257

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

258

async fn resume_and_fork_append_permissions_messages() -> Result<()> {

259

skip_if_no_network!(Ok(()));

260

261

let server = start_mock_server().await;

262

let _req1 = mount_sse_once(&server, sse_completed("resp-1")).await;

263

let req2 = mount_sse_once(&server, sse_completed("resp-2")).await;

264

let req3 = mount_sse_once(&server, sse_completed("resp-3")).await;

265

let req4 = mount_sse_once(&server, sse_completed("resp-4")).await;

266

267

let mut builder = test_codex().with_config(|config| {

268

config.approval_policy = Constrained::allow_any(AskForApproval::OnRequest);

269

});

270

let initial = builder.build(&server).await?;

271

let rollout_path = initial.session_configured.rollout_path.clone();

272

let home = initial.home.clone();

273

274

initial

275

.codex

276

.submit(Op::UserInput {

277

items: vec![UserInput::Text {

278

text: "hello 1".into(),

279

}],

280

final_output_json_schema: None,

281

})

282

.await?;

283

wait_for_event(&initial.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;

284

285

initial

286

.codex

287

.submit(Op::OverrideTurnContext {

288

cwd: None,

289

approval_policy: Some(AskForApproval::Never),

290

sandbox_policy: None,

291

model: None,

292

effort: None,

293

summary: None,

294

})

295

.await?;

296

297

initial

298

.codex

299

.submit(Op::UserInput {

300

items: vec![UserInput::Text {

301

text: "hello 2".into(),

302

}],

303

final_output_json_schema: None,

304

})

305

.await?;

306

wait_for_event(&initial.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;

307

308

let body2 = req2.single_request().body_json();

309

let input2 = body2["input"].as_array().expect("input array");

310

let permissions_base = permissions_texts(input2);

311

assert_eq!(permissions_base.len(), 2);

312

313

builder = builder.with_config(|config| {

314

config.approval_policy = Constrained::allow_any(AskForApproval::UnlessTrusted);

315

});

316

let resumed = builder.resume(&server, home, rollout_path.clone()).await?;

317

resumed

318

.codex

319

.submit(Op::UserInput {

320

items: vec![UserInput::Text {

321

text: "after resume".into(),

322

}],

323

final_output_json_schema: None,

324

})

325

.await?;

326

wait_for_event(&resumed.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;

327

328

let body3 = req3.single_request().body_json();

329

let input3 = body3["input"].as_array().expect("input array");

330

let permissions_resume = permissions_texts(input3);

331

assert_eq!(permissions_resume.len(), permissions_base.len() + 1);

332

assert_eq!(

333

&permissions_resume[..permissions_base.len()],

334

permissions_base.as_slice()

335

);

336

assert!(!permissions_base.contains(permissions_resume.last().expect("new permissions")));

337

338

let mut fork_config = initial.config.clone();

339

fork_config.approval_policy = Constrained::allow_any(AskForApproval::UnlessTrusted);

340

let forked = initial

341

.thread_manager

342

.fork_thread(usize::MAX, fork_config, rollout_path)

343

.await?;

344

forked

345

.thread

346

.submit(Op::UserInput {

347

items: vec![UserInput::Text {

348

text: "after fork".into(),

349

}],

350

final_output_json_schema: None,

351

})

352

.await?;

353

wait_for_event(&forked.thread, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;

354

355

let body4 = req4.single_request().body_json();

356

let input4 = body4["input"].as_array().expect("input array");

357

let permissions_fork = permissions_texts(input4);

358

assert_eq!(permissions_fork.len(), permissions_base.len() + 2);

359

assert_eq!(

360

&permissions_fork[..permissions_base.len()],

361

permissions_base.as_slice()

362

);

363

let new_permissions = &permissions_fork[permissions_base.len()..];

364

assert_eq!(new_permissions.len(), 2);

365

assert_eq!(new_permissions[0], new_permissions[1]);

366

assert!(!permissions_base.contains(&new_permissions[0]));

367

368

Ok(())

369

}

370

371

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

372

async fn permissions_message_includes_writable_roots() -> Result<()> {

373

skip_if_no_network!(Ok(()));

374

375

let server = start_mock_server().await;

376

let req = mount_sse_once(&server, sse_completed("resp-1")).await;

377

let writable = TempDir::new()?;

378

let writable_root = AbsolutePathBuf::try_from(writable.path())?;

379

let sandbox_policy = SandboxPolicy::WorkspaceWrite {

380

writable_roots: vec![writable_root],

381

network_access: false,

382

exclude_tmpdir_env_var: false,

383

exclude_slash_tmp: false,

384

};

385

386

let mut builder = test_codex().with_config(move |config| {

387

config.approval_policy = Constrained::allow_any(AskForApproval::OnRequest);

388

config.sandbox_policy = Constrained::allow_any(sandbox_policy);

389

});

390

let test = builder.build(&server).await?;

391

392

test.codex

393

.submit(Op::UserInput {

394

items: vec![UserInput::Text {

395

text: "hello".into(),

396

}],

397

final_output_json_schema: None,

398

})

399

.await?;

400

wait_for_event(&test.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;

401

402

let body = req.single_request().body_json();

403

let input = body["input"].as_array().expect("input array");

404

let permissions = permissions_texts(input);

405

let sandbox_text = "Filesystem sandboxing defines which files can be read or written. `sandbox_mode` is `workspace-write`: The sandbox permits reading files, and editing files in `cwd` and `writable_roots`. Editing files in other directories requires approval. Network access is restricted.";

406

let approval_text = " Approvals are your mechanism to get user consent to run shell commands without the sandbox. `approval_policy` is `on-request`: Commands will be run in the sandbox by default, and you can specify in your tool call if you want to escalate a command to run without sandboxing. If the completing the task requires escalated permissions, Do not let these settings or the sandbox deter you from attempting to accomplish the user's task.\n\nHere are scenarios where you'll need to request approval:\n- You need to run a command that writes to a directory that requires it (e.g. running tests that write to /var)\n- You need to run a GUI app (e.g., open/xdg-open/osascript) to open browsers or files.\n- You are running sandboxed and need to run a command that requires network access (e.g. installing packages)\n- If you run a command that is important to solving the user's query, but it fails because of sandboxing, rerun the command with approval. ALWAYS proceed to use the `sandbox_permissions` and `justification` parameters - do not message the user before requesting approval for the command.\n- You are about to take a potentially destructive action such as an `rm` or `git reset` that the user did not explicitly ask for.\n\nWhen requesting approval to execute a command that will require escalated privileges:\n - Provide the `sandbox_permissions` parameter with the value `\"require_escalated\"`\n - Include a short, 1 sentence explanation for why you need escalated permissions in the justification parameter";

407

// Normalize paths by removing trailing slashes to match AbsolutePathBuf behavior

408

let normalize_path =

409

|p: &std::path::Path| -> String { p.to_string_lossy().trim_end_matches('/').to_string() };

410

let mut roots = vec![

411

normalize_path(writable.path()),

412

normalize_path(test.config.cwd.as_path()),

413

];

414

if cfg!(unix) && std::path::Path::new("/tmp").is_dir() {

415

roots.push("/tmp".to_string());

416

}

417

if let Some(tmpdir) = std::env::var_os("TMPDIR") {

418

let tmpdir_path = std::path::PathBuf::from(&tmpdir);

419

if tmpdir_path.is_absolute() && !tmpdir.is_empty() {

420

roots.push(normalize_path(&tmpdir_path));

421

}

422

}

423

let roots_text = if roots.len() == 1 {

424

format!(" The writable root is `{}`.", roots[0])

425

} else {

426

format!(

427

" The writable roots are {}.",

428

roots

429

.iter()

430

.map(|root| format!("`{root}`"))

431

.collect::<Vec<_>>()

432

.join(", ")

433

)

434

};

435

let expected = format!(

436

"<permissions instructions>{sandbox_text}{approval_text}{roots_text}</permissions instructions>"

437

);

438

// Normalize line endings to handle Windows vs Unix differences

439

let normalize_line_endings = |s: &str| s.replace("\r\n", "\n");

440

let expected_normalized = normalize_line_endings(&expected);

441

let actual_normalized: Vec<String> = permissions

442

.iter()

443

.map(|s| normalize_line_endings(s))

444

.collect();

445

assert_eq!(actual_normalized, vec![expected_normalized]);

446

447

Ok(())

448

}

6Test Adjustments for New Message Structure

0 / 9

send_message.rscodex-rs/app-server/tests/suite

+28

Mark as viewed

1

use anyhow::Result;

1

use anyhow::Result;

5 linesAll 13 lines5 lines

5 linesAll 13 lines5 lines

15

use codex_protocol::models::ContentItem;

15

use codex_protocol::models::ContentItem;

16

use codex_protocol::models::DeveloperInstructions;

16

use codex_protocol::models::ResponseItem;

17

use codex_protocol::models::ResponseItem;

18

use codex_protocol::protocol::AskForApproval;

17

use codex_protocol::protocol::RawResponseItemEvent;

19

use codex_protocol::protocol::RawResponseItemEvent;

20

use codex_protocol::protocol::SandboxPolicy;

18

use core_test_support::responses;

21

use core_test_support::responses;

19

use pretty_assertions::assert_eq;

22

use pretty_assertions::assert_eq;

20

use std::path::Path;

23

use std::path::Path;

24

use std::path::PathBuf;

5 linesAll 122 lines5 lines

5 linesAll 122 lines5 lines

143

#[tokio::test]

147

#[tokio::test]

144

async fn test_send_message_raw_notifications_opt_in() -> Result<()> {

148

async fn test_send_message_raw_notifications_opt_in() -> Result<()> {

5 linesAll 43 lines5 lines

5 linesAll 43 lines5 lines

188

let send_id = mcp

192

let send_id = mcp

189

.send_send_user_message_request(SendUserMessageParams {

193

.send_send_user_message_request(SendUserMessageParams {

All 5 lines

All 5 lines

195

.await?;

199

.await?;

196

200

201

let permissions = read_raw_response_item(&mut mcp, conversation_id).await;

202

assert_permissions_message(&permissions);

203

197

let developer = read_raw_response_item(&mut mcp, conversation_id).await;

204

let developer = read_raw_response_item(&mut mcp, conversation_id).await;

5 linesAll 28 lines5 lines

5 linesAll 28 lines5 lines

226

}

233

}

5 linesAll 116 lines5 lines

5 linesAll 116 lines5 lines

1

Copy + paste350–369

350

fn assert_permissions_message(item: &ResponseItem) {

351

match item {

352

ResponseItem::Message { role, content, .. } => {

353

assert_eq!(role, "developer");

354

let texts = content_texts(content);

355

let expected = DeveloperInstructions::from_policy(

356

&SandboxPolicy::DangerFullAccess,

357

AskForApproval::Never,

358

&PathBuf::from("/tmp"),

359

)

360

.into_text();

361

assert_eq!(

362

texts,

363

vec![expected.as_str()],

364

"expected permissions developer message, got {texts:?}"

365

);

366

}

367

other => panic!("expected permissions message, got {other:?}"),

368

}

369

}

370

5 linesAll 64 lines5 lines

5 linesAll 64 lines5 lines

truncation.rscodex-rs/core/src/rollout

+1

Mark as viewed

5 linesAll 70 lines5 lines

5 linesAll 70 lines5 lines

71

#[cfg(test)]

71

#[cfg(test)]

72

mod tests {

72

mod tests {

5 linesAll 116 lines5 lines

5 linesAll 116 lines5 lines

189

#[tokio::test]

189

#[tokio::test]

190

async fn ignores_session_prefix_messages_when_truncating_rollout_from_start() {

190

async fn ignores_session_prefix_messages_when_truncating_rollout_from_start() {

5 linesAll 13 lines5 lines

5 linesAll 13 lines5 lines

204

let truncated = truncate_rollout_before_nth_user_message_from_start(&rollout_items, 1);

204

let truncated = truncate_rollout_before_nth_user_message_from_start(&rollout_items, 1);

205

let expected: Vec<RolloutItem> = vec![

205

let expected: Vec<RolloutItem> = vec![

206

RolloutItem::ResponseItem(items[0].clone()),

206

RolloutItem::ResponseItem(items[0].clone()),

207

RolloutItem::ResponseItem(items[1].clone()),

207

RolloutItem::ResponseItem(items[1].clone()),

208

RolloutItem::ResponseItem(items[2].clone()),

208

RolloutItem::ResponseItem(items[2].clone()),

209

RolloutItem::ResponseItem(items[3].clone()),

209

];

210

];

210

211

211

assert_eq!(

212

assert_eq!(

212

serde_json::to_value(&truncated).unwrap(),

213

serde_json::to_value(&truncated).unwrap(),

213

serde_json::to_value(&expected).unwrap()

214

serde_json::to_value(&expected).unwrap()

214

);

215

);

215

}

216

}

216

}

217

}

thread_manager.rscodex-rs/core/src

+1

Mark as viewed

5 linesAll 302 lines5 lines

5 linesAll 302 lines5 lines

303

#[cfg(test)]

303

#[cfg(test)]

304

mod tests {

304

mod tests {

5 linesAll 78 lines5 lines

5 linesAll 78 lines5 lines

383

#[tokio::test]

383

#[tokio::test]

384

async fn ignores_session_prefix_messages_when_truncating() {

384

async fn ignores_session_prefix_messages_when_truncating() {

5 linesAll 16 lines5 lines

5 linesAll 16 lines5 lines

401

let expected: Vec<RolloutItem> = vec![

401

let expected: Vec<RolloutItem> = vec![

402

RolloutItem::ResponseItem(items[0].clone()),

402

RolloutItem::ResponseItem(items[0].clone()),

403

RolloutItem::ResponseItem(items[1].clone()),

403

RolloutItem::ResponseItem(items[1].clone()),

404

RolloutItem::ResponseItem(items[2].clone()),

404

RolloutItem::ResponseItem(items[2].clone()),

PA2

CommentR405

405

RolloutItem::ResponseItem(items[3].clone()),

405

];

406

];

All 5 lines

All 5 lines

411

}

412

}

412

}

413

}

client.rscodex-rs/core/tests/suite

+88−36

Mark as viewed

5 linesAll 157 lines5 lines

5 linesAll 157 lines5 lines

158

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

158

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

159

async fn resume_includes_initial_messages_and_sends_prior_items() {

159

async fn resume_includes_initial_messages_and_sends_prior_items() {

5 linesAll 118 lines5 lines

5 linesAll 118 lines5 lines

278

// 1) Assert initial_messages only includes existing EventMsg entries; response items are not converted

278

// 1) Assert initial_messages only includes existing EventMsg entries; response items are not converted

All 6 lines

All 6 lines

285

assert_eq!(initial_json, expected_initial_json);

285

assert_eq!(initial_json, expected_initial_json);

286

286

287

// 2) Submit new input; the request body must include the prior item followed by the new user input.

287

// 2) Submit new input; the request body must include the prior items, then initial context, then new user input.

288

codex

288

codex

289

.submit(Op::UserInput {

289

.submit(Op::UserInput {

All 4 lines

All 4 lines

294

})

294

})

295

.await

295

.await

296

.unwrap();

296

.unwrap();

297

wait_for_event(&codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;

297

wait_for_event(&codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;

298

298

299

let request = resp_mock.single_request();

299

let request = resp_mock.single_request();

300

let request_body = request.body_json();

300

let request_body = request.body_json();

301

let expected_input = json!([

301

let input = request_body["input"].as_array().expect("input array");

302

{

302

let messages: Vec<(String, String)> = input

303

"type": "message",

303

.iter()

304

"role": "user",

304

.filter_map(|item| {

305

"content": [{ "type": "input_text", "text": "resumed user message" }]

305

let role = item.get("role")?.as_str()?;

306

},

306

let text = item

307

{

307

.get("content")?

308

"type": "message",

308

.as_array()?

309

"role": "assistant",

309

.first()?

310

"content": [{ "type": "output_text", "text": "resumed assistant message" }]

310

.get("text")?

311

},

311

.as_str()?;

312

{

312

Some((role.to_string(), text.to_string()))

313

"type": "message",

313

})

314

"role": "user",

314

.collect();

315

"content": [{ "type": "input_text", "text": "hello" }]

315

let pos_prior_user = messages

316

}

316

.iter()

317

]);

317

.position(|(role, text)| role == "user" && text == "resumed user message")

318

assert_eq!(request_body["input"], expected_input);

318

.expect("prior user message");

319

let pos_prior_assistant = messages

320

.iter()

321

.position(|(role, text)| role == "assistant" && text == "resumed assistant message")

322

.expect("prior assistant message");

323

let pos_permissions = messages

324

.iter()

325

.position(|(role, text)| role == "developer" && text.contains("`approval_policy`"))

326

.expect("permissions message");

327

let pos_user_instructions = messages

328

.iter()

329

.position(|(role, text)| {

330

role == "user"

331

&& text.contains("be nice")

332

&& (text.starts_with("# AGENTS.md instructions for ")

333

|| text.starts_with("<user_instructions>"))

334

})

335

.expect("user instructions");

336

let pos_environment = messages

337

.iter()

338

.position(|(role, text)| role == "user" && text.contains("<environment_context>"))

339

.expect("environment context");

340

let pos_new_user = messages

341

.iter()

342

.position(|(role, text)| role == "user" && text == "hello")

343

.expect("new user message");

344

345

assert!(pos_prior_user < pos_prior_assistant);

346

assert!(pos_prior_assistant < pos_permissions);

347

assert!(pos_permissions < pos_user_instructions);

348

assert!(pos_user_instructions < pos_environment);

349

assert!(pos_environment < pos_new_user);

319

}

350

}

5 linesAll 252 lines5 lines

5 linesAll 252 lines5 lines

572

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

603

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

573

async fn includes_user_instructions_message_in_request() {

604

async fn includes_user_instructions_message_in_request() {

5 linesAll 38 lines5 lines

5 linesAll 38 lines5 lines

612

let request = resp_mock.single_request();

643

let request = resp_mock.single_request();

613

let request_body = request.body_json();

644

let request_body = request.body_json();

614

645

615

assert!(

646

assert!(

All 4 lines

All 4 lines

620

);

651

);

621

assert_message_role(&request_body["input"][0], "user");

652

assert_message_role(&request_body["input"][0], "developer");

622

assert_message_starts_with(&request_body["input"][0], "# AGENTS.md instructions for ");

653

let permissions_text = request_body["input"][0]["content"][0]["text"]

623

assert_message_ends_with(&request_body["input"][0], "</INSTRUCTIONS>");

654

.as_str()

624

let ui_text = request_body["input"][0]["content"][0]["text"]

655

.expect("invalid permissions message content");

656

assert!(

657

permissions_text.contains("`sandbox_mode`"),

658

"expected permissions message to mention sandbox_mode, got {permissions_text:?}"

659

);

660

661

assert_message_role(&request_body["input"][1], "user");

662

assert_message_starts_with(&request_body["input"][1], "# AGENTS.md instructions for ");

663

assert_message_ends_with(&request_body["input"][1], "</INSTRUCTIONS>");

664

let ui_text = request_body["input"][1]["content"][0]["text"]

625

.as_str()

665

.as_str()

626

.expect("invalid message content");

666

.expect("invalid message content");

627

assert!(ui_text.contains("<INSTRUCTIONS>"));

667

assert!(ui_text.contains("<INSTRUCTIONS>"));

628

assert!(ui_text.contains("be nice"));

668

assert!(ui_text.contains("be nice"));

629

assert_message_role(&request_body["input"][1], "user");

669

assert_message_role(&request_body["input"][2], "user");

630

assert_message_starts_with(&request_body["input"][1], "<environment_context>");

670

assert_message_starts_with(&request_body["input"][2], "<environment_context>");

631

assert_message_ends_with(&request_body["input"][1], "</environment_context>");

671

assert_message_ends_with(&request_body["input"][2], "</environment_context>");

632

}

672

}

633

673

634

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

674

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

635

async fn skills_append_to_instructions() {

675

async fn skills_append_to_instructions() {

5 linesAll 46 lines5 lines

5 linesAll 46 lines5 lines

682

let request = resp_mock.single_request();

722

let request = resp_mock.single_request();

683

let request_body = request.body_json();

723

let request_body = request.body_json();

684

724

685

assert_message_role(&request_body["input"][0], "user");

725

assert_message_role(&request_body["input"][0], "developer");

686

let instructions_text = request_body["input"][0]["content"][0]["text"]

726

727

assert_message_role(&request_body["input"][1], "user");

728

let instructions_text = request_body["input"][1]["content"][0]["text"]

687

.as_str()

729

.as_str()

5 linesAll 15 lines5 lines

5 linesAll 15 lines5 lines

703

let _codex_home_guard = codex_home;

745

let _codex_home_guard = codex_home;

704

}

746

}

5 linesAll 303 lines5 lines

5 linesAll 303 lines5 lines

1008

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

1050

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

1009

async fn includes_developer_instructions_message_in_request() {

1051

async fn includes_developer_instructions_message_in_request() {

5 linesAll 39 lines5 lines

5 linesAll 39 lines5 lines

1049

let request = resp_mock.single_request();

1091

let request = resp_mock.single_request();

1050

let request_body = request.body_json();

1092

let request_body = request.body_json();

1051

1093

1094

let permissions_text = request_body["input"][0]["content"][0]["text"]

1095

.as_str()

1096

.expect("invalid permissions message content");

1097

1052

assert!(

1098

assert!(

All 4 lines

All 4 lines

1057

);

1103

);

1058

assert_message_role(&request_body["input"][0], "developer");

1104

assert_message_role(&request_body["input"][0], "developer");

1059

assert_message_equals(&request_body["input"][0], "be useful");

1105

assert!(

1060

assert_message_role(&request_body["input"][1], "user");

1106

permissions_text.contains("`sandbox_mode`"),

1061

assert_message_starts_with(&request_body["input"][1], "# AGENTS.md instructions for ");

1107

"expected permissions message to mention sandbox_mode, got {permissions_text:?}"

1062

assert_message_ends_with(&request_body["input"][1], "</INSTRUCTIONS>");

1108

);

1063

let ui_text = request_body["input"][1]["content"][0]["text"]

1109

1110

assert_message_role(&request_body["input"][1], "developer");

1111

assert_message_equals(&request_body["input"][1], "be useful");

1112

assert_message_role(&request_body["input"][2], "user");

1113

assert_message_starts_with(&request_body["input"][2], "# AGENTS.md instructions for ");

1114

assert_message_ends_with(&request_body["input"][2], "</INSTRUCTIONS>");

1115

let ui_text = request_body["input"][2]["content"][0]["text"]

1064

.as_str()

1116

.as_str()

1065

.expect("invalid message content");

1117

.expect("invalid message content");

1066

assert!(ui_text.contains("<INSTRUCTIONS>"));

1118

assert!(ui_text.contains("<INSTRUCTIONS>"));

1067

assert!(ui_text.contains("be nice"));

1119

assert!(ui_text.contains("be nice"));

1068

assert_message_role(&request_body["input"][2], "user");

1120

assert_message_role(&request_body["input"][3], "user");

1069

assert_message_starts_with(&request_body["input"][2], "<environment_context>");

1121

assert_message_starts_with(&request_body["input"][3], "<environment_context>");

1070

assert_message_ends_with(&request_body["input"][2], "</environment_context>");

1122

assert_message_ends_with(&request_body["input"][3], "</environment_context>");

1071

}

1123

}

5 linesAll 799 lines5 lines

5 linesAll 799 lines5 lines

compact.rscodex-rs/core/tests/suite

+12−4

Mark as viewed

5 linesAll 459 lines5 lines

5 linesAll 459 lines5 lines

460

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

460

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

461

async fn multiple_auto_compact_per_task_runs_after_token_limit_hit() {

461

async fn multiple_auto_compact_per_task_runs_after_token_limit_hit() {

5 linesAll 126 lines5 lines

5 linesAll 126 lines5 lines

588

fn normalize_inputs(values: &[serde_json::Value]) -> Vec<serde_json::Value> {

588

fn normalize_inputs(values: &[serde_json::Value]) -> Vec<serde_json::Value> {

589

values

589

values

590

.iter()

590

.iter()

591

.filter(|value| {

591

.filter(|value| {

All 8 lines

All 8 lines

600

let text = value

600

let text = value

All 4 lines

All 4 lines

605

.and_then(|text| text.as_str());

605

.and_then(|text| text.as_str());

606

606

607

// Ignore the cached UI prefix (project docs + skills) since it is not relevant to

607

// Ignore cached prefix messages (project docs + permissions) since they are not

608

// compaction behavior and can change as bundled skills evolve.

608

// relevant to compaction behavior and can change as bundled prompts evolve.

609

let role = value.get("role").and_then(|role| role.as_str());

610

if role == Some("developer")

611

&& text.is_some_and(|text| text.contains("`sandbox_mode`"))

612

{

613

return false;

614

}

609

!text.is_some_and(|text| text.starts_with("# AGENTS.md instructions for "))

615

!text.is_some_and(|text| text.starts_with("# AGENTS.md instructions for "))

610

})

616

})

611

.cloned()

617

.cloned()

612

.collect()

618

.collect()

613

}

619

}

5 linesAll 383 lines5 lines

5 linesAll 383 lines5 lines

997

}

1003

}

5 linesAll 561 lines5 lines

5 linesAll 561 lines5 lines

1559

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

1565

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

1560

async fn manual_compact_twice_preserves_latest_user_messages() {

1566

async fn manual_compact_twice_preserves_latest_user_messages() {

5 linesAll 161 lines5 lines

5 linesAll 161 lines5 lines

1722

let mut final_output = requests

1728

let mut final_output = requests

1723

.last()

1729

.last()

1724

.unwrap_or_else(|| panic!("final turn request missing for {final_user_message}"))

1730

.unwrap_or_else(|| panic!("final turn request missing for {final_user_message}"))

1725

.input()

1731

.input()

1726

.into_iter()

1732

.into_iter()

1727

.collect::<VecDeque<_>>();

1733

.collect::<VecDeque<_>>();

1728

1734

1729

// System prompt

1735

// Permissions developer message

1736

final_output.pop_front();

1737

// User instructions (project docs/skills)

1730

final_output.pop_front();

1738

final_output.pop_front();

1731

// Developer instructions

1739

// Environment context

1732

final_output.pop_front();

1740

final_output.pop_front();

5 linesAll 41 lines5 lines

5 linesAll 41 lines5 lines

1774

}

1782

}

5 linesAll 344 lines5 lines

5 linesAll 344 lines5 lines

compact_resume_fork.rscodex-rs/core/tests/suite

+96−4

Mark as viewed

5 linesAll 139 lines5 lines

5 linesAll 139 lines5 lines

140

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

140

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

141

/// Scenario: compact an initial conversation, resume it, fork one turn back, and

141

/// Scenario: compact an initial conversation, resume it, fork one turn back, and

142

/// ensure the model-visible history matches expectations at each request.

142

/// ensure the model-visible history matches expectations at each request.

143

async fn compact_resume_and_fork_preserve_model_history_view() {

143

async fn compact_resume_and_fork_preserve_model_history_view() {

5 linesAll 67 lines5 lines

5 linesAll 67 lines5 lines

211

let expected_model = requests[0]["model"]

211

let expected_model = requests[0]["model"]

212

.as_str()

212

.as_str()

213

.unwrap_or_default()

213

.unwrap_or_default()

214

.to_string();

214

.to_string();

215

let prompt = requests[0]["instructions"]

215

let prompt = requests[0]["instructions"]

216

.as_str()

216

.as_str()

217

.unwrap_or_default()

217

.unwrap_or_default()

218

.to_string();

218

.to_string();

219

let user_instructions = requests[0]["input"][0]["content"][0]["text"]

219

let permissions_message = requests[0]["input"][0].clone();

220

let user_instructions = requests[0]["input"][1]["content"][0]["text"]

220

.as_str()

221

.as_str()

221

.unwrap_or_default()

222

.unwrap_or_default()

222

.to_string();

223

.to_string();

223

let environment_context = requests[0]["input"][1]["content"][0]["text"]

224

let environment_context = requests[0]["input"][2]["content"][0]["text"]

224

.as_str()

225

.as_str()

5 linesAll 13 lines5 lines

5 linesAll 13 lines5 lines

238

let summary_after_fork = extract_summary_message(&requests[4], SUMMARY_TEXT);

239

let summary_after_fork = extract_summary_message(&requests[4], SUMMARY_TEXT);

239

let user_turn_1 = json!(

240

let user_turn_1 = json!(

240

{

241

{

241

"model": expected_model,

242

"model": expected_model,

242

"instructions": prompt,

243

"instructions": prompt,

243

"input": [

244

"input": [

245

permissions_message,

244

{

246

{

245

"type": "message",

247

"type": "message",

5 linesAll 27 lines5 lines

5 linesAll 27 lines5 lines

273

}

275

}

274

],

276

],

5 linesAll 12 lines5 lines

5 linesAll 12 lines5 lines

287

});

289

});

288

let compact_1 = json!(

290

let compact_1 = json!(

289

{

291

{

290

"model": expected_model,

292

"model": expected_model,

291

"instructions": prompt,

293

"instructions": prompt,

292

"input": [

294

"input": [

295

permissions_message,

293

{

296

{

294

"type": "message",

297

"type": "message",

5 linesAll 47 lines5 lines

5 linesAll 47 lines5 lines

342

}

345

}

343

],

346

],

5 linesAll 12 lines5 lines

5 linesAll 12 lines5 lines

356

});

359

});

357

let user_turn_2_after_compact = json!(

360

let user_turn_2_after_compact = json!(

358

{

361

{

359

"model": expected_model,

362

"model": expected_model,

360

"instructions": prompt,

363

"instructions": prompt,

361

"input": [

364

"input": [

365

permissions_message,

362

{

366

{

363

"type": "message",

367

"type": "message",

5 linesAll 38 lines5 lines

5 linesAll 38 lines5 lines

402

}

406

}

403

],

407

],

5 linesAll 12 lines5 lines

5 linesAll 12 lines5 lines

416

});

420

});

417

let usert_turn_3_after_resume = json!(

421

let usert_turn_3_after_resume = json!(

418

{

422

{

419

"model": expected_model,

423

"model": expected_model,

420

"instructions": prompt,

424

"instructions": prompt,

421

"input": [

425

"input": [

426

permissions_message,

422

{

427

{

423

"type": "message",

428

"type": "message",

5 linesAll 47 lines5 lines

5 linesAll 47 lines5 lines

471

]

476

]

472

},

477

},

478

permissions_message,

479

{

480

"type": "message",

481

"role": "user",

482

"content": [

483

{

484

"type": "input_text",

485

"text": user_instructions

486

}

487

]

488

},

489

{

490

"type": "message",

491

"role": "user",

492

"content": [

493

{

494

"type": "input_text",

495

"text": environment_context

496

}

497

]

498

},

473

{

499

{

474

"type": "message",

500

"type": "message",

All 7 lines

All 7 lines

482

}

508

}

483

],

509

],

5 linesAll 12 lines5 lines

5 linesAll 12 lines5 lines

496

});

522

});

497

let user_turn_3_after_fork = json!(

523

let user_turn_3_after_fork = json!(

498

{

524

{

499

"model": expected_model,

525

"model": expected_model,

500

"instructions": prompt,

526

"instructions": prompt,

501

"input": [

527

"input": [

528

permissions_message,

502

{

529

{

503

"type": "message",

530

"type": "message",

5 linesAll 47 lines5 lines

5 linesAll 47 lines5 lines

551

]

578

]

552

},

579

},

580

permissions_message,

581

{

582

"type": "message",

583

"role": "user",

584

"content": [

585

{

586

"type": "input_text",

587

"text": user_instructions

588

}

589

]

590

},

591

{

592

"type": "message",

593

"role": "user",

594

"content": [

595

{

596

"type": "input_text",

597

"text": environment_context

598

}

599

]

600

},

601

permissions_message,

602

{

603

"type": "message",

604

"role": "user",

605

"content": [

606

{

607

"type": "input_text",

608

"text": user_instructions

609

}

610

]

611

},

612

{

613

"type": "message",

614

"role": "user",

615

"content": [

616

{

617

"type": "input_text",

618

"text": environment_context

619

}

620

]

621

},

553

{

622

{

554

"type": "message",

623

"type": "message",

All 7 lines

All 7 lines

562

}

631

}

563

],

632

],

5 linesAll 12 lines5 lines

5 linesAll 12 lines5 lines

576

});

645

});

5 linesAll 13 lines5 lines

5 linesAll 13 lines5 lines

590

}

659

}

591

660

592

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

661

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

593

/// Scenario: after the forked branch is compacted, resuming again should reuse

662

/// Scenario: after the forked branch is compacted, resuming again should reuse

594

/// the compacted history and only append the new user message.

663

/// the compacted history and only append the new user message.

595

async fn compact_resume_after_second_compaction_preserves_history() {

664

async fn compact_resume_after_second_compaction_preserves_history() {

5 linesAll 66 lines5 lines

5 linesAll 66 lines5 lines

662

// hard coded test

731

// hard coded test

663

let prompt = requests[0]["instructions"]

732

let prompt = requests[0]["instructions"]

664

.as_str()

733

.as_str()

665

.unwrap_or_default()

734

.unwrap_or_default()

666

.to_string();

735

.to_string();

667

let user_instructions = requests[0]["input"][0]["content"][0]["text"]

736

let permissions_message = requests[0]["input"][0].clone();

737

let user_instructions = requests[0]["input"][1]["content"][0]["text"]

668

.as_str()

738

.as_str()

669

.unwrap_or_default()

739

.unwrap_or_default()

670

.to_string();

740

.to_string();

671

let environment_instructions = requests[0]["input"][1]["content"][0]["text"]

741

let environment_instructions = requests[0]["input"][2]["content"][0]["text"]

672

.as_str()

742

.as_str()

All 8 lines

All 8 lines

681

let mut expected = json!([

751

let mut expected = json!([

682

{

752

{

683

"instructions": prompt,

753

"instructions": prompt,

684

"input": [

754

"input": [

755

permissions_message,

685

{

756

{

686

"type": "message",

757

"type": "message",

5 linesAll 37 lines5 lines

5 linesAll 37 lines5 lines

724

]

795

]

725

},

796

},

797

permissions_message,

798

{

799

"type": "message",

800

"role": "user",

801

"content": [

802

{

803

"type": "input_text",

804

"text": user_instructions

805

}

806

]

807

},

808

{

809

"type": "message",

810

"role": "user",

811

"content": [

812

{

813

"type": "input_text",

814

"text": environment_instructions

815

}

816

]

817

},

726

{

818

{

727

"type": "message",

819

"type": "message",

All 7 lines

All 7 lines

735

}

827

}

736

],

828

],

737

}

829

}

738

]);

830

]);

5 linesAll 12 lines5 lines

5 linesAll 12 lines5 lines

751

}

843

}

5 linesAll 194 lines5 lines

5 linesAll 194 lines5 lines

fork_thread.rscodex-rs/core/tests/suite

+4−2

Mark as viewed

5 linesAll 27 lines5 lines

5 linesAll 27 lines5 lines

28

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

28

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

29

async fn fork_thread_twice_drops_to_first_message() {

29

async fn fork_thread_twice_drops_to_first_message() {

5 linesAll 109 lines5 lines

5 linesAll 109 lines5 lines

139

// GetHistory on fork1 flushed; the file is ready.

139

// GetHistory on fork1 flushed; the file is ready.

140

let fork1_items = read_items(&fork1_path);

140

let fork1_items = read_items(&fork1_path);

141

assert!(fork1_items.len() > expected_after_first.len());

141

pretty_assertions::assert_eq!(

142

pretty_assertions::assert_eq!(

142

serde_json::to_value(&fork1_items).unwrap(),

143

serde_json::to_value(&fork1_items[..expected_after_first.len()]).unwrap(),

143

serde_json::to_value(&expected_after_first).unwrap()

144

serde_json::to_value(&expected_after_first).unwrap()

144

);

145

);

5 linesAll 18 lines5 lines

5 linesAll 18 lines5 lines

163

let expected_after_second: Vec<RolloutItem> = fork1_items[..cut_last_on_fork1].to_vec();

164

let expected_after_second: Vec<RolloutItem> = fork1_items[..cut_last_on_fork1].to_vec();

164

let fork2_items = read_items(&fork2_path);

165

let fork2_items = read_items(&fork2_path);

166

assert!(fork2_items.len() > expected_after_second.len());

165

pretty_assertions::assert_eq!(

167

pretty_assertions::assert_eq!(

166

serde_json::to_value(&fork2_items).unwrap(),

168

serde_json::to_value(&fork2_items[..expected_after_second.len()]).unwrap(),

167

serde_json::to_value(&expected_after_second).unwrap()

169

serde_json::to_value(&expected_after_second).unwrap()

168

);

170

);

169

}

171

}

prompt_caching.rscodex-rs/core/tests/suite

+85−82

Mark as viewed

5 linesAll 33 lines5 lines

5 linesAll 33 lines5 lines

34

fn default_env_context_str(cwd: &str, shell: &Shell) -> String {

34

fn default_env_context_str(cwd: &str, shell: &Shell) -> String {

35

let shell_name = shell.name();

35

let shell_name = shell.name();

36

format!(

36

format!(

37

r#"<environment_context>

37

r#"<environment_context>

38

<cwd>{cwd}</cwd>

38

<cwd>{cwd}</cwd>

39

<approval_policy>on-request</approval_policy>

40

<sandbox_mode>read-only</sandbox_mode>

41

<network_access>restricted</network_access>

42

<shell>{shell_name}</shell>

39

<shell>{shell_name}</shell>

43

</environment_context>"#

40

</environment_context>"#

44

)

41

)

45

}

42

}

5 linesAll 170 lines5 lines

5 linesAll 170 lines5 lines

216

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

213

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

217

async fn prefixes_context_and_instructions_once_and_consistently_across_requests()

214

async fn prefixes_context_and_instructions_once_and_consistently_across_requests()

218

-> anyhow::Result<()> {

215

-> anyhow::Result<()> {

5 linesAll 34 lines5 lines

5 linesAll 34 lines5 lines

253

let body1 = req1.single_request().body_json();

250

let body1 = req1.single_request().body_json();

254

let input1 = body1["input"].as_array().expect("input array");

251

let input1 = body1["input"].as_array().expect("input array");

255

assert_eq!(input1.len(), 3, "expected cached prefix + env + user msg");

252

assert_eq!(

253

input1.len(),

254

4,

InformationalR255-260

255

"expected permissions + cached prefix + env + user msg"

256

);

256

257

257

let ui_text = input1[0]["content"][0]["text"]

258

let ui_text = input1[1]["content"][0]["text"]

All 7 lines

All 7 lines

265

let shell = default_user_shell();

266

let shell = default_user_shell();

266

let cwd_str = config.cwd.to_string_lossy();

267

let cwd_str = config.cwd.to_string_lossy();

267

let expected_env_text = default_env_context_str(&cwd_str, &shell);

268

let expected_env_text = default_env_context_str(&cwd_str, &shell);

268

assert_eq!(

269

assert_eq!(

269

input1[1],

270

input1[2],

270

text_user_input(expected_env_text),

271

text_user_input(expected_env_text),

271

"expected environment context after UI message"

272

"expected environment context after UI message"

272

);

273

);

273

assert_eq!(input1[2], text_user_input("hello 1".to_string()));

274

assert_eq!(input1[3], text_user_input("hello 1".to_string()));

5 linesAll 10 lines5 lines

5 linesAll 10 lines5 lines

284

Ok(())

285

Ok(())

285

}

286

}

286

287

287

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

288

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

288

async fn overrides_turn_context_but_keeps_cached_prefix_and_key_constant() -> anyhow::Result<()> {

289

async fn overrides_turn_context_but_keeps_cached_prefix_and_key_constant() -> anyhow::Result<()> {

5 linesAll 25 lines5 lines

5 linesAll 25 lines5 lines

314

let writable = TempDir::new().unwrap();

315

let writable = TempDir::new().unwrap();

316

let new_policy = SandboxPolicy::WorkspaceWrite {

317

writable_roots: vec![writable.path().try_into().unwrap()],

318

network_access: true,

319

exclude_tmpdir_env_var: true,

320

exclude_slash_tmp: true,

321

};

315

codex

322

codex

316

.submit(Op::OverrideTurnContext {

323

.submit(Op::OverrideTurnContext {

317

cwd: None,

324

cwd: None,

318

approval_policy: Some(AskForApproval::Never),

325

approval_policy: Some(AskForApproval::Never),

319

sandbox_policy: Some(SandboxPolicy::WorkspaceWrite {

326

sandbox_policy: Some(new_policy.clone()),

320

writable_roots: vec![writable.path().try_into().unwrap()],

321

network_access: true,

322

exclude_tmpdir_env_var: true,

323

exclude_slash_tmp: true,

324

}),

325

model: Some("o3".to_string()),

327

model: Some("o3".to_string()),

326

effort: Some(Some(ReasoningEffort::High)),

328

effort: Some(Some(ReasoningEffort::High)),

327

summary: Some(ReasoningSummary::Detailed),

329

summary: Some(ReasoningSummary::Detailed),

328

})

330

})

329

.await?;

331

.await?;

5 linesAll 22 lines5 lines

5 linesAll 22 lines5 lines

352

let expected_user_message_2 = serde_json::json!({

354

let expected_user_message_2 = serde_json::json!({

353

"type": "message",

355

"type": "message",

354

"role": "user",

356

"role": "user",

355

"content": [ { "type": "input_text", "text": "hello 2" } ]

357

"content": [ { "type": "input_text", "text": "hello 2" } ]

356

});

358

});

357

// After overriding the turn context, the environment context should be emitted again

359

let expected_permissions_msg = body1["input"][0].clone();

358

// reflecting the new approval policy and sandbox settings. Omit cwd because it did

360

// After overriding the turn context, emit a new permissions message.

359

// not change.

361

let body1_input = body1["input"].as_array().expect("input array");

360

let shell = default_user_shell();

362

let expected_permissions_msg_2 = body2["input"][body1_input.len()].clone();

361

let expected_env_text_2 = format!(

363

assert_ne!(

362

r#"<environment_context>

364

expected_permissions_msg_2, expected_permissions_msg,

363

<approval_policy>never</approval_policy>

365

"expected updated permissions message after override"

364

<sandbox_mode>workspace-write</sandbox_mode>

365

<network_access>enabled</network_access>

366

<writable_roots>

367

<root>{}</root>

368

</writable_roots>

369

<shell>{}</shell>

370

</environment_context>"#,

371

writable.path().display(),

372

shell.name()

373

);

366

);

374

let expected_env_msg_2 = serde_json::json!({

367

let mut expected_body2 = body1["input"].as_array().expect("input array").to_vec();

375

"type": "message",

368

expected_body2.push(expected_permissions_msg_2);

376

"role": "user",

369

expected_body2.push(expected_user_message_2);

377

"content": [ { "type": "input_text", "text": expected_env_text_2 } ]

370

assert_eq!(body2["input"], serde_json::Value::Array(expected_body2));

378

});

379

let expected_body2 = serde_json::json!(

380

[

381

body1["input"].as_array().unwrap().as_slice(),

382

[expected_env_msg_2, expected_user_message_2].as_slice(),

383

]

384

.concat()

385

);

386

assert_eq!(body2["input"], expected_body2);

387

371

388

Ok(())

372

Ok(())

389

}

373

}

390

374

391

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

375

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

392

async fn override_before_first_turn_emits_environment_context() -> anyhow::Result<()> {

376

async fn override_before_first_turn_emits_environment_context() -> anyhow::Result<()> {

5 linesAll 48 lines5 lines

5 linesAll 48 lines5 lines

441

assert!(

425

assert!(

442

env_texts

426

!env_texts.is_empty(),

443

.iter()

427

"expected environment context to be emitted: {env_texts:?}"

444

.any(|text| text.contains("<approval_policy>never</approval_policy>")),

445

"environment context should reflect overridden approval policy: {env_texts:?}"

446

);

428

);

447

429

448

let env_count = input

430

let env_count = input

449

.iter()

431

.iter()

450

.filter(|msg| {

432

.filter(|msg| {

5 linesAll 12 lines5 lines

5 linesAll 12 lines5 lines

463

})

445

})

464

.count();

446

.count();

465

assert_eq!(

447

assert!(

466

env_count, 2,

448

env_count >= 1,

467

"environment context should appear exactly twice, found {env_count}"

449

"environment context should appear at least once, found {env_count}"

450

);

451

452

let permissions_texts: Vec<&str> = input

453

.iter()

454

.filter_map(|msg| {

455

let role = msg["role"].as_str()?;

456

if role != "developer" {

457

return None;

458

}

459

msg["content"]

460

.as_array()

461

.and_then(|content| content.first())

462

.and_then(|item| item["text"].as_str())

463

})

464

.collect();

465

assert!(

466

permissions_texts

467

.iter()

468

.any(|text| text.contains("`approval_policy` is `never`")),

469

"permissions message should reflect overridden approval policy: {permissions_texts:?}"

468

);

470

);

5 linesAll 16 lines5 lines

5 linesAll 16 lines5 lines

485

}

487

}

486

488

487

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

489

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

488

async fn per_turn_overrides_keep_cached_prefix_and_key_constant() -> anyhow::Result<()> {

490

async fn per_turn_overrides_keep_cached_prefix_and_key_constant() -> anyhow::Result<()> {

5 linesAll 27 lines5 lines

5 linesAll 27 lines5 lines

516

let writable = TempDir::new().unwrap();

518

let writable = TempDir::new().unwrap();

519

let new_policy = SandboxPolicy::WorkspaceWrite {

520

writable_roots: vec![AbsolutePathBuf::try_from(writable.path()).unwrap()],

521

network_access: true,

522

exclude_tmpdir_env_var: true,

523

exclude_slash_tmp: true,

524

};

517

codex

525

codex

518

.submit(Op::UserTurn {

526

.submit(Op::UserTurn {

All 5 lines

All 5 lines

524

sandbox_policy: SandboxPolicy::WorkspaceWrite {

532

sandbox_policy: new_policy.clone(),

525

writable_roots: vec![AbsolutePathBuf::try_from(writable.path()).unwrap()],

526

network_access: true,

527

exclude_tmpdir_env_var: true,

528

exclude_slash_tmp: true,

529

},

All 4 lines

All 4 lines

534

})

537

})

535

.await?;

538

.await?;

5 linesAll 20 lines5 lines

5 linesAll 20 lines5 lines

556

let expected_env_text_2 = format!(

559

let expected_env_text_2 = format!(

557

r#"<environment_context>

560

r#"<environment_context>

558

<cwd>{}</cwd>

561

<cwd>{}</cwd>

559

<approval_policy>never</approval_policy>

560

<sandbox_mode>workspace-write</sandbox_mode>

561

<network_access>enabled</network_access>

562

<writable_roots>

563

<root>{}</root>

564

</writable_roots>

565

<shell>{}</shell>

562

<shell>{}</shell>

566

</environment_context>"#,

563

</environment_context>"#,

567

new_cwd.path().display(),

564

new_cwd.path().display(),

568

writable.path().display(),

565

shell.name()

569

shell.name(),

570

);

566

);

All 5 lines

All 5 lines

576

let expected_body2 = serde_json::json!(

572

let expected_permissions_msg = body1["input"][0].clone();

577

[

573

let body1_input = body1["input"].as_array().expect("input array");

578

body1["input"].as_array().unwrap().as_slice(),

574

let expected_permissions_msg_2 = body2["input"][body1_input.len() + 1].clone();

579

[expected_env_msg_2, expected_user_message_2].as_slice(),

575

assert_ne!(

580

]

576

expected_permissions_msg_2, expected_permissions_msg,

581

.concat()

577

"expected updated permissions message after per-turn override"

582

);

578

);

583

assert_eq!(body2["input"], expected_body2);

579

let mut expected_body2 = body1_input.to_vec();

580

expected_body2.push(expected_env_msg_2);

581

expected_body2.push(expected_permissions_msg_2);

582

expected_body2.push(expected_user_message_2);

583

assert_eq!(body2["input"], serde_json::Value::Array(expected_body2));

584

584

585

Ok(())

585

Ok(())

586

}

586

}

587

587

588

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

588

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

589

async fn send_user_turn_with_no_changes_does_not_send_environment_context() -> anyhow::Result<()> {

589

async fn send_user_turn_with_no_changes_does_not_send_environment_context() -> anyhow::Result<()> {

5 linesAll 61 lines5 lines

5 linesAll 61 lines5 lines

651

let expected_ui_msg = body1["input"][0].clone();

651

let expected_permissions_msg = body1["input"][0].clone();

652

let expected_ui_msg = body1["input"][1].clone();

All 7 lines

All 7 lines

659

let expected_input_1 = serde_json::Value::Array(vec![

660

let expected_input_1 = serde_json::Value::Array(vec![

661

expected_permissions_msg.clone(),

660

expected_ui_msg.clone(),

662

expected_ui_msg.clone(),

661

expected_env_msg_1.clone(),

663

expected_env_msg_1.clone(),

662

expected_user_message_1.clone(),

664

expected_user_message_1.clone(),

663

]);

665

]);

664

assert_eq!(body1["input"], expected_input_1);

666

assert_eq!(body1["input"], expected_input_1);

665

667

666

let expected_user_message_2 = text_user_input("hello 2".to_string());

668

let expected_user_message_2 = text_user_input("hello 2".to_string());

667

let expected_input_2 = serde_json::Value::Array(vec![

669

let expected_input_2 = serde_json::Value::Array(vec![

670

expected_permissions_msg,

668

expected_ui_msg,

671

expected_ui_msg,

669

expected_env_msg_1,

672

expected_env_msg_1,

670

expected_user_message_1,

673

expected_user_message_1,

671

expected_user_message_2,

674

expected_user_message_2,

672

]);

675

]);

673

assert_eq!(body2["input"], expected_input_2);

676

assert_eq!(body2["input"], expected_input_2);

674

677

675

Ok(())

678

Ok(())

676

}

679

}

677

680

678

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

681

#[tokio::test(flavor = "multi_thread", worker_threads = 2)]

679

async fn send_user_turn_with_changes_sends_environment_context() -> anyhow::Result<()> {

682

async fn send_user_turn_with_changes_sends_environment_context() -> anyhow::Result<()> {

5 linesAll 61 lines5 lines

5 linesAll 61 lines5 lines

741

let expected_ui_msg = body1["input"][0].clone();

744

let expected_permissions_msg = body1["input"][0].clone();

745

let expected_ui_msg = body1["input"][1].clone();

All 5 lines

All 5 lines

747

let expected_input_1 = serde_json::Value::Array(vec![

751

let expected_input_1 = serde_json::Value::Array(vec![

752

expected_permissions_msg.clone(),

748

expected_ui_msg.clone(),

753

expected_ui_msg.clone(),

749

expected_env_msg_1.clone(),

754

expected_env_msg_1.clone(),

750

expected_user_message_1.clone(),

755

expected_user_message_1.clone(),

751

]);

756

]);

752

assert_eq!(body1["input"], expected_input_1);

757

assert_eq!(body1["input"], expected_input_1);

753

758

754

let shell_name = shell.name();

759

let body1_input = body1["input"].as_array().expect("input array");

755

let expected_env_msg_2 = text_user_input(format!(

760

let expected_permissions_msg_2 = body2["input"][body1_input.len()].clone();

756

r#"<environment_context>

761

assert_ne!(

757

<approval_policy>never</approval_policy>

762

expected_permissions_msg_2, expected_permissions_msg,

758

<sandbox_mode>danger-full-access</sandbox_mode>

763

"expected updated permissions message after policy change"

759

<network_access>enabled</network_access>

764

);

760

<shell>{shell_name}</shell>

761

</environment_context>"#

762

));

763

let expected_user_message_2 = text_user_input("hello 2".to_string());

765

let expected_user_message_2 = text_user_input("hello 2".to_string());

764

let expected_input_2 = serde_json::Value::Array(vec![

766

let expected_input_2 = serde_json::Value::Array(vec![

767

expected_permissions_msg,

765

expected_ui_msg,

768

expected_ui_msg,

766

expected_env_msg_1,

769

expected_env_msg_1,

767

expected_user_message_1,

770

expected_user_message_1,

768

expected_env_msg_2,

771

expected_permissions_msg_2,

769

expected_user_message_2,

772

expected_user_message_2,

770

]);

773

]);

771

assert_eq!(body2["input"], expected_input_2);

774

assert_eq!(body2["input"], expected_input_2);

772

775

773

Ok(())

776

Ok(())

774

}

777

}

codex_tool.rscodex-rs/mcp-server/tests/suite

+17−14

Mark as viewed

5 linesAll 332 lines5 lines

5 linesAll 332 lines5 lines

333

async fn codex_tool_passes_base_instructions() -> anyhow::Result<()> {

333

async fn codex_tool_passes_base_instructions() -> anyhow::Result<()> {

5 linesAll 45 lines5 lines

5 linesAll 45 lines5 lines

379

let requests = server.received_requests().await.unwrap();

379

let requests = server.received_requests().await.unwrap();

380

let request = requests[0].body_json::<serde_json::Value>()?;

380

let request = requests[0].body_json::<serde_json::Value>()?;

381

let instructions = request["messages"][0]["content"].as_str().unwrap();

381

let instructions = request["messages"][0]["content"].as_str().unwrap();

382

assert!(instructions.starts_with("You are a helpful assistant."));

382

assert!(instructions.starts_with("You are a helpful assistant."));

383

383

384

let developer_msg = request["messages"]

384

let developer_messages: Vec<&serde_json::Value> = request["messages"]

385

.as_array()

385

.as_array()

386

.and_then(|messages| {

386

.unwrap()

387

messages

387

.iter()

388

.iter()

388

.filter(|msg| msg.get("role").and_then(|role| role.as_str()) == Some("developer"))

389

.find(|msg| msg.get("role").and_then(|role| role.as_str()) == Some("developer"))

389

.collect();

390

})

390

let developer_contents: Vec<&str> = developer_messages

391

.unwrap();

391

.iter()

392

let developer_content = developer_msg

392

.filter_map(|msg| msg.get("content").and_then(|value| value.as_str()))

393

.get("content")

393

.collect();

394

.and_then(|value| value.as_str())

394

assert!(

395

.unwrap();

395

developer_contents

396

.iter()

397

.any(|content| content.contains("`sandbox_mode`")),

398

"expected permissions developer message, got {developer_contents:?}"

399

);

396

assert!(

400

assert!(

397

!developer_content.contains('<'),

401

developer_contents.contains(&"Foreshadow upcoming tool calls."),

398

"expected developer instructions without XML tags, got `{developer_content}`"

402

"expected developer instructions in developer messages, got {developer_contents:?}"

399

);

403

);

400

assert_eq!(developer_content, "Foreshadow upcoming tool calls.");

401

404

402

Ok(())

405

Ok(())

403

}

406

}

5 linesAll 87 lines5 lines

5 linesAll 87 lines5 lines

End of changesChat about this PR

InfoChat

1 Potential bug

Permissions message not updated when cwd changes in WorkspaceWrite mode

Bugcodex.rs:1017

5 Flags

Resume/fork intentionally duplicates initial context for policy synchronization

codex.rs:856-859

EnvironmentContext simplified - sandbox info moved to permissions message

environment_context.rs:12-16

Prompt templates use placeholder interpolation for network_access

models.rs:272-281

Test expectations updated for new message ordering

prompt_caching.rs:255-260

Static prompts removed from markdown - no longer bundled in instructions

prompt.md

ChecksPassed33/33

All checks passed

Reviewers3

DYdylan-hurd-oai

CHchatgpt-codex-connector

PApakrym-oai

Assignees

No assignees

Labels

No labels assigned

Introducing Devin Review

Devin Review is an intelligent PR review tool that helps you better understand your team's code before you ship it.

1

Intelligently organizes your diff

Groups changes into logical sections with clear explanations

→

2

Detects moved or copied code

Keeps the diff clean by highlighting copy-pasted or moved code

3

Analyzes your PR

Highlights potential bugs and flags areas that deserve extra careful review

Ask questions or request edits in the chat.

Ask Devin anything about this PR (Ctrl+I)

Authorization Response