Home
Wiki
Review
Docs
PR #8961
1New DeveloperInstructions Builder
10 files+237−28
user_instructions.rscodex-rs/core/src
BUILD.bazelcodex-rs/protocol
models.rscodex-rs/protocol/src
never.mdcodex-rs/protocol/src/prompts/permissions/approval_policy
on_failure.mdcodex-rs/protocol/src/prompts/permissions/approval_policy
+5 more files
Read explanation
The DeveloperInstructions struct moved from core/src/user_instructions.rs to protocol/src/models.rs and gained a new factory method that constructs permissions messages from runtime configuration.
models.rs:189-231
rust
impl DeveloperInstructions { pub fn new<T: Into<String>>(text: T) -> Self pub fn into_text(self) -> String pub fn concat(self, other: impl Into<Self>) -> Self /// Main entry point: builds permissions /// message from current policy settings pub fn from_policy( sandbox_policy: &SandboxPolicy, approval_policy: AskForApproval, cwd: &Path, ) -> Self { // Derive network_access from sandbox_policy // Derive sandbox_mode and writable_roots // Combine into structured message }}
The message content comes from template files that get include_str!'d at compile time:
prompts/permissions/├── approval_policy/│ ├── never.md│ ├── on_failure.md│ ├── on_request.md│ └── unless_trusted.md└── sandbox_mode/ ├── danger_full_access.md ├── read_only.md └── workspace_write.md
Each template contains the prose for that mode. For example, workspace_write.mdworkspace_write.md:1:
markdown
Filesystem sandboxing defines which filescan be read or written. `sandbox_mode` is`workspace-write`: The sandbox permitsreading files, and editing files in `cwd`and `writable_roots`. Editing files inother directories requires approval.Network access is {network_access}.
The {network_access} placeholder is replaced with enabled or restricted at runtime models.rs:272-281.
2Permissions Message Injection in Session
1 file+51−7
codex.rscodex-rs/core/src
Read explanation
The permissions message is injected at three points in the session lifecycle.
At session start — added to build_initial_context():
codex.rs:1364-1374
rust
fn build_initial_context(&self, ctx: &TurnContext) -> Vec<ResponseItem>{ let mut items = Vec::with_capacity(4); // was 3 // NEW: permissions message first items.push( DeveloperInstructions::from_policy( &ctx.sandbox_policy, ctx.approval_policy, &ctx.cwd, ).into() ); // Then: developer_instructions, // user_instructions, environment ...}
When policies change mid-session — new method build_permissions_update_item():
codex.rs:1012-1032
rust
fn build_permissions_update_item( &self, previous: Option<&Arc<TurnContext>>, next: &TurnContext,) -> Option<ResponseItem> { let prev = previous?; if prev.sandbox_policy == next.sandbox_policy && prev.approval_policy == next.approval_policy { return None; // no change, no message } Some(DeveloperInstructions::from_policy(...).into())}
This is called during user input handling codex.rs:1895-1901 alongside the existing environment update check.
On resume/fork — initial context appended after reconstructed history: codex.rs:855-859
rust
// After reconstructing history from rollout:let initial_context = self.build_initial_context(&turn_context);self.record_conversation_items( &turn_context, &initial_context).await;
3Simplified EnvironmentContext
1 file+20−223
environment_context.rscodex-rs/core/src
Read explanation
Since sandbox/approval info moved to the permissions message, EnvironmentContext was slimmed down significantly.
environment_context.rs:17-25
diff
pub struct EnvironmentContext { pub cwd: Option<PathBuf>,- pub approval_policy: Option<AskForApproval>,- pub sandbox_mode: Option<SandboxMode>,- pub network_access: Option<NetworkAccess>,- pub writable_roots: Option<Vec<AbsolutePathBuf>>, pub shell: Shell, }
The constructor simplified from a complex match on SandboxPolicy to:
environment_context.rs:19-20
rust
pub fn new(cwd: Option<PathBuf>, shell: Shell) -> Self { Self { cwd, shell }}
The equals_except_shell() comparison now only checks cwdenvironment_context.rs:27-34, and the XML serialization no longer emits approval/sandbox/network tags environment_context.rs:51-70.
4Removed Static Prompt Sections
7 files−255
gpt-5.1-codex-max_prompt.mdcodex-rs/core
gpt-5.2-codex_prompt.mdcodex-rs/core
gpt_5_1_prompt.mdcodex-rs/core
gpt_5_2_prompt.mdcodex-rs/core
gpt_5_codex_prompt.mdcodex-rs/core
+2 more files
Read explanation
The "Codex CLI harness, sandboxing, and approvals" section (~35 lines) was removed from all prompt markdown files since this content is now generated dynamically:
5New permissions_messages Test Suite
2 files+449
mod.rscodex-rs/core/tests/suite
permissions_messages.rscodex-rs/core/tests/suite
Read explanation
A comprehensive test file validates the new permissions messaging behavior permissions_messages.rs:1-448:
permissions_message_sent_once_on_start— verifies single emission at session startpermissions_message_added_on_override_change— verifies re-emission when policy changespermissions_message_not_added_when_no_change— verifies no duplicate when policy unchangedresume_replays_permissions_messages— verifies history replay on resumeresume_and_fork_append_permissions_messages— verifies fresh context appended on resume/forkpermissions_message_includes_writable_roots— verifies writable roots formatting
6Test Adjustments for New Message Structure
9 files+332−142
send_message.rscodex-rs/app-server/tests/suite
truncation.rscodex-rs/core/src/rollout
thread_manager.rscodex-rs/core/src
client.rscodex-rs/core/tests/suite
compact.rscodex-rs/core/tests/suite
+4 more files
Read explanation
Tests throughout the codebase were updated to account for the new permissions message at input[0]. Most changes are mechanical index shifts (e.g., input[0] → input[1]) and additions of permissions_message to expected JSON structures.
Key patterns:
- Request body assertions now check for permissions message first client.rs:652-659
- Resume/fork tests verify message ordering client.rs:345-349
- Compact tests filter out permissions messages when comparing compact.rs:607-614
- Fork tests account for initial context appended after truncation fork_thread.rs:141-143
Merged
Assemble sandbox/approval/network prompts dynamically
main dynamic/permissions/instructions
30 files
+1089−655
DescriptionDiscussion41Commits48
Devin's AI analysis
This PR moves sandbox/approval/network permission instructions from static system prompt files into dynamically generated developer messages. Previously, these instructions were baked into markdown files and sent identically to every session. Now they're constructed at runtime based on actual configuration.
Key change: A new DeveloperInstructions::from_policy() method models.rs:204-231 builds a developer-role message from the current SandboxPolicy and AskForApproval settings. This message is:
- Injected at session start via
build_initial_context()codex.rs:1367-1374 - Re-injected when policies change mid-session via
build_permissions_update_item()codex.rs:1012-1032 - Appended after reconstructed history on resume/fork codex.rs:855-859
Consequence: EnvironmentContext no longer carries approval_policy, sandbox_mode, network_access, or writable_roots fields—these moved to the permissions message. The struct now only holds cwd and shell.
Request structure change:
Before: input[0]: user instructions (AGENTS.md) input[1]: environment context input[2]: user messageAfter: input[0]: permissions (developer role) input[1]: user instructions (AGENTS.md) input[2]: environment context input[3]: user message
- Add a single builder for developer permissions messaging that accepts SandboxPolicy and approval policy. This builder now drives the developer “permissions” message that’s injected at session start and any time sandbox/approval settings change.
- Trim EnvironmentContext to only include cwd, writable roots, and shell; removed sandbox/approval/network duplication and adjusted XML serialization and tests accordingly.
Follow-up: adding a config value to replace the developer permissions message for custom sandboxes.
Filter
AI
aibrahim-oai112d
@codex review
CH
chatgpt-codex-connector
commented
112d
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit:a40e5a755a
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
models.rsOutdatedResolved
AI
aibrahim-oai112d
@codex review this and make sure tests changes didn't invalidate tests original intentions
CH
chatgpt-codex-connector
commented
112d
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit:2240249141
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
prompt_caching.rsOutdatedResolved
AI
aibrahim-oai112d
@codex review this
CH
chatgpt-codex-connector
commented
112d
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit:fc5c68019b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
on_request.mdOutdatedResolved
AI
aibrahim-oai112d
@codex review
CH
chatgpt-codex-connector112d
Codex Review: Didn't find any major issues. Swish!
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
PApakrym-oai
commented
112d
models.rsL175
PA
pakrym-oai112d
a bit strange to see all of this formatting stuff in the protocol
AI
aibrahim-oai112d
Yeah, just so that they can be methods on the structs. I can move them with the structs to core.
DY
dylan-hurd-oai109d
Agree w/ pavel. Is the plan to start using DeveloperInstructions as a serialized struct?
AI
aibrahim-oai109d
no I can move it
PApakrym-oai
commented
112d
models.rsOutdated
PA
pakrym-oai112d
A lot of methods. Do we have this many callsites that we need all the overrides?
AI
aibrahim-oai112d
Not really, we can have one method but just wanted to divide responsibility so we don't have a super long function with complicated logic.
PA
pakrym-oai112d
You can return strings from helper method and make them private.
PApakrym-oai
commented
112d
models.rsOutdated
PA
pakrym-oai112d
Why are these values encoded here? is there an existing place where we convert SandboxPolicy to SandboxMode/network?
PApakrym-oai
commented
112d
models.rsL289Resolved
PApakrym-oai
commented
112d
thread_manager.rsL405
PA
pakrym-oai112d
is this expected?
👍1
AI
aibrahim-oai109d
yes we added a new message
PApakrym-oai
commented
112d
environment_context.rsL14
PA
pakrym-oai112d
we left so little here I wonder whether it matters.
AI
aibrahim-oai112d
I want to move it to dev slowly per the plan.
AI
aibrahim-oai111d
@codex do we append a new permission message on resume and fork next to the older messages? If yes, add an integration test. If not, change behavior and add integration test.
CH
chatgpt-codex-connector111d
Summary
- Appended a fresh permissions developer message when resuming or forking sessions and persisted it alongside reconstructed history for the new turn context. codex-rs/core/src/codex.rsL836-L861
- Updated resume/fork history expectations and expanded permissions integration coverage to include a forked-session case and the extra permissions message count. codex-rs/core/src/codex.rsL2961-L3079 codex-rs/core/tests/suite/permissions_messages.rsL182-L327
Testing
- ✅
/root/.cargo/bin/just fmt - ✅
/root/.cargo/bin/just fix -p codex-core - ❌
cargo test -p codex-core(failed:default_client::tests::test_create_client_sets_default_headers;exec::tests::kill_child_process_group_kills_grandchildren_on_timeout;shell_snapshot::tests::linux_sh_snapshot_includes_sections;models_manager::manager::tests::refresh_available_models_{drops_removed_remote_models,refetches_when_cache_stale,sorts_and_marks_default,uses_cache_when_fresh};unified_exec::tests::{reusing_completed_process_returns_unknown_process,multi_unified_exec_sessions})
DY
dylan-hurd-oai
commented
109d
Some small comments but mostly LGTM!
codex.rsL1032
DY
dylan-hurd-oai109d
(non-blocking) Should we move this out of codex.rs? Either into environment_context.rs or a new file
AI
aibrahim-oai109d
it's on turn context unfortunately
codex.rsL2982Resolved
1,744 lines left
1New DeveloperInstructions Builder
0 / 10
user_instructions.rscodex-rs/core/src
−28
Mark as viewed
5 linesAll 77 lines5 lines
5 linesAll 77 lines5 lines
1
Move78–104
78
#[derive(Debug, Clone, Serialize, Deserialize, PartialEq)]
79
#[serde(rename = "developer_instructions", rename_all = "snake_case")]
80
pub(crate) struct DeveloperInstructions {
81
text: String,
82
}
83
84
impl DeveloperInstructions {
85
pub fn new<T: Into<String>>(text: T) -> Self {
86
Self { text: text.into() }
87
}
88
89
pub fn into_text(self) -> String {
90
self.text
91
}
92
}
93
94
impl From<DeveloperInstructions> for ResponseItem {
95
fn from(di: DeveloperInstructions) -> Self {
96
ResponseItem::Message {
97
id: None,
98
role: "developer".to_string(),
99
content: vec![ContentItem::InputText {
100 text: di.into_text(),
101 }],
102
}
103
}
104
}
105
5 linesAll 88 lines5 lines
5 linesAll 88 lines5 lines
BUILD.bazelcodex-rs/protocol
+1
Mark as viewed
1
load("//:defs.bzl", "codex_rust_crate")
1
load("//:defs.bzl", "codex_rust_crate")
2
``
2
``
3
codex_rust_crate(
3
codex_rust_crate(
4
name = "protocol",
4
name = "protocol",
5
crate_name = "codex_protocol",
5
crate_name = "codex_protocol",
6
compile_data = glob(["src/prompts/permissions/**/*.md"]),
6
)
7
)
models.rscodex-rs/protocol/src
+218
Mark as viewed
1
use std::collections::HashMap;
1
use std::collections::HashMap;
2
use std::path::Path;
5 linesAll 10 lines5 lines
5 linesAll 10 lines5 lines
13
use crate::config_types::SandboxMode;
14
use crate::protocol::AskForApproval;
15
use crate::protocol::NetworkAccess;
16
use crate::protocol::SandboxPolicy;
17
use crate::protocol::WritableRoot;
5 linesAll 149 lines5 lines
5 linesAll 149 lines5 lines
1
Move167–294
167
/// Developer-provided guidance that is injected into a turn as a developer role
168
/// message.
169
#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, JsonSchema, TS)]
170
#[serde(rename = "developer_instructions", rename_all = "snake_case")]
171
pub struct DeveloperInstructions {
172
text: String,
173
}
174
PA4
CommentR175
175
const APPROVAL_POLICY_NEVER: &str = include_str!("prompts/permissions/approval_policy/never.md");
176
const APPROVAL_POLICY_UNLESS_TRUSTED: &str =
177
include_str!("prompts/permissions/approval_policy/unless_trusted.md");
178
const APPROVAL_POLICY_ON_FAILURE: &str =
179
include_str!("prompts/permissions/approval_policy/on_failure.md");
180
const APPROVAL_POLICY_ON_REQUEST: &str =
181
include_str!("prompts/permissions/approval_policy/on_request.md");
182
183
const SANDBOX_MODE_DANGER_FULL_ACCESS: &str =
184
include_str!("prompts/permissions/sandbox_mode/danger_full_access.md");
185
const SANDBOX_MODE_WORKSPACE_WRITE: &str =
186
include_str!("prompts/permissions/sandbox_mode/workspace_write.md");
187
const SANDBOX_MODE_READ_ONLY: &str = include_str!("prompts/permissions/sandbox_mode/read_only.md");
188
189
impl DeveloperInstructions {
190
pub fn new<T: Into<String>>(text: T) -> Self {
191
Self { text: text.into() }
192
}
193
194
pub fn into_text(self) -> String {
195
self.text
196
}
197
198
pub fn concat(self, other: impl Into<DeveloperInstructions>) -> Self {
199
let mut text = self.text;
200
text.push_str(&other.into().text);
201
Self { text }
202
}
203
204
pub fn from_policy(
205
sandbox_policy: &SandboxPolicy,
206
approval_policy: AskForApproval,
207
cwd: &Path,
208
) -> Self {
209
let network_access = if sandbox_policy.has_full_network_access() {
210
NetworkAccess::Enabled
211
} else {
212
NetworkAccess::Restricted
213
};
214
215
let (sandbox_mode, writable_roots) = match sandbox_policy {
216
SandboxPolicy::DangerFullAccess => (SandboxMode::DangerFullAccess, None),
217
SandboxPolicy::ReadOnly => (SandboxMode::ReadOnly, None),
218
SandboxPolicy::ExternalSandbox { .. } => (SandboxMode::DangerFullAccess, None),
219
SandboxPolicy::WorkspaceWrite { .. } => {
220
let roots = sandbox_policy.get_writable_roots_with_cwd(cwd);
221
(SandboxMode::WorkspaceWrite, Some(roots))
222
}
223
};
224
225
DeveloperInstructions::from_permissions_with_network(
226
sandbox_mode,
227
network_access,
228
approval_policy,
229
writable_roots,
230
)
231
}
232
233
fn from_permissions_with_network(
234
sandbox_mode: SandboxMode,
235
network_access: NetworkAccess,
236
approval_policy: AskForApproval,
237
writable_roots: Option<Vec<WritableRoot>>,
238
) -> Self {
239
let start_tag = DeveloperInstructions::new("<permissions instructions>");
240
let end_tag = DeveloperInstructions::new("</permissions instructions>");
241
start_tag
242
.concat(DeveloperInstructions::sandbox_text(
243
sandbox_mode,
244
network_access,
245
))
246
.concat(DeveloperInstructions::from(approval_policy))
247
.concat(DeveloperInstructions::from_writable_roots(writable_roots))
248
.concat(end_tag)
249
}
250
251
fn from_writable_roots(writable_roots: Option<Vec<WritableRoot>>) -> Self {
252
let Some(roots) = writable_roots else {
253
return DeveloperInstructions::new("");
254
};
255
256
if roots.is_empty() {
257
return DeveloperInstructions::new("");
258
}
259
260
let roots_list: Vec<String> = roots
261
.iter()
262
.map(|r| format!("`{}`", r.root.to_string_lossy()))
263
.collect();
264
let text = if roots_list.len() == 1 {
265
format!(" The writable root is {}.", roots_list[0])
266
} else {
267
format!(" The writable roots are {}.", roots_list.join(", "))
268
};
269
DeveloperInstructions::new(text)
270
}
271
InformationalR272-281
272
fn sandbox_text(mode: SandboxMode, network_access: NetworkAccess) -> DeveloperInstructions {
273
let template = match mode {
274
SandboxMode::DangerFullAccess => SANDBOX_MODE_DANGER_FULL_ACCESS.trim_end(),
275
SandboxMode::WorkspaceWrite => SANDBOX_MODE_WORKSPACE_WRITE.trim_end(),
276
SandboxMode::ReadOnly => SANDBOX_MODE_READ_ONLY.trim_end(),
277
};
278
let text = template.replace("{network_access}", &network_access.to_string());
279
280
DeveloperInstructions::new(text)
281
}
282
}
283
284
impl From<DeveloperInstructions> for ResponseItem {
285
fn from(di: DeveloperInstructions) -> Self {
286
ResponseItem::Message {
287
id: None,
288
role: "developer".to_string(),
PA
CommentR289
Resolved
289
content: vec![ContentItem::InputText {
290 text: di.into_text(),
291 }],
292
}
293
}
294
}
295
296
impl From<SandboxMode> for DeveloperInstructions {
297
fn from(mode: SandboxMode) -> Self {
298
let network_access = match mode {
299
SandboxMode::DangerFullAccess => NetworkAccess::Enabled,
300
SandboxMode::WorkspaceWrite | SandboxMode::ReadOnly => NetworkAccess::Restricted,
301
};
302
303
DeveloperInstructions::sandbox_text(mode, network_access)
304
}
305
}
306
307
impl From<AskForApproval> for DeveloperInstructions {
308
fn from(mode: AskForApproval) -> Self {
309
let text = match mode {
310
AskForApproval::Never => APPROVAL_POLICY_NEVER.trim_end(),
311
AskForApproval::UnlessTrusted => APPROVAL_POLICY_UNLESS_TRUSTED.trim_end(),
312
AskForApproval::OnFailure => APPROVAL_POLICY_ON_FAILURE.trim_end(),
313
AskForApproval::OnRequest => APPROVAL_POLICY_ON_REQUEST.trim_end(),
314
};
315
316
DeveloperInstructions::new(text)
317
}
318
}
319
5 linesAll 389 lines5 lines
5 linesAll 389 lines5 lines
550
#[cfg(test)]
709
#[cfg(test)]
551
mod tests {
710
mod tests {
552
use super::*;
711
use super::*;
712
use crate::config_types::SandboxMode;
713
use crate::protocol::AskForApproval;
All 4 lines
All 4 lines
718
use std::path::PathBuf;
557
use tempfile::tempdir;
719
use tempfile::tempdir;
558
720
721
#[test]
722
fn converts_sandbox_mode_into_developer_instructions() {
723
assert_eq!(
724
DeveloperInstructions::from(SandboxMode::WorkspaceWrite),
725
DeveloperInstructions::new(
726
"Filesystem sandboxing defines which files can be read or written. `sandbox_mode` is `workspace-write`: The sandbox permits reading files, and editing files in `cwd` and `writable_roots`. Editing files in other directories requires approval. Network access is restricted."
727
)
728
);
729
730
assert_eq!(
731
DeveloperInstructions::from(SandboxMode::ReadOnly),
732
DeveloperInstructions::new(
733
"Filesystem sandboxing defines which files can be read or written. `sandbox_mode` is `read-only`: The sandbox only permits reading files. Network access is restricted."
734
)
735
);
736
}
737
738
#[test]
739
fn builds_permissions_with_network_access_override() {
740
let instructions = DeveloperInstructions::from_permissions_with_network(
741
SandboxMode::WorkspaceWrite,
742
NetworkAccess::Enabled,
743
AskForApproval::OnRequest,
744
None,
745
);
746
747
let text = instructions.into_text();
748
assert!(
749
text.contains("Network access is enabled."),
750
"expected network access to be enabled in message"
751
);
752
assert!(
753
text.contains("`approval_policy` is `on-request`"),
754
"expected approval guidance to be included"
755
);
756
}
757
758
#[test]
759
fn builds_permissions_from_policy() {
760
let policy = SandboxPolicy::WorkspaceWrite {
761
writable_roots: vec![],
762
network_access: true,
763
exclude_tmpdir_env_var: false,
764
exclude_slash_tmp: false,
765
};
766
767
let instructions = DeveloperInstructions::from_policy(
768
&policy,
769
AskForApproval::UnlessTrusted,
770
&PathBuf::from("/tmp"),
771
);
772
let text = instructions.into_text();
773
assert!(text.contains("Network access is enabled."));
774
assert!(text.contains("`approval_policy` is `unless-trusted`"));
775
}
776
5 linesAll 311 lines5 lines
5 linesAll 311 lines5 lines
870
}
1088
}
never.mdcodex-rs/protocol/src/prompts/permissions/approval_policy
+1
AddedMark as viewed
1
Approvals are your mechanism to get user consent to run shell commands without the sandbox. `approval_policy` is `never`: This is a non-interactive mode where you may NEVER ask the user for approval to run commands. Instead, you must always persist and work around constraints to solve the task for the user. You MUST do your utmost best to finish the task and validate your work before yielding. If this mode is paired with `danger-full-access`, take advantage of it to deliver the best outcome for the user. Further, in this mode, your default testing philosophy is overridden: Even if you don't see local patterns for testing, you may add tests and scripts to validate your work. Just remove them before yielding.
on_failure.mdcodex-rs/protocol/src/prompts/permissions/approval_policy
+1
AddedMark as viewed
1
Approvals are your mechanism to get user consent to run shell commands without the sandbox. `approval_policy` is `on-failure`: The harness will allow all commands to run in the sandbox (if enabled), and failures will be escalated to the user for approval to run again without the sandbox.
on_request.mdcodex-rs/protocol/src/prompts/permissions/approval_policy
+12
AddedMark as viewed
1
Move1–12
1
Approvals are your mechanism to get user consent to run shell commands without the sandbox. `approval_policy` is `on-request`: Commands will be run in the sandbox by default, and you can specify in your tool call if you want to escalate a command to run without sandboxing. If the completing the task requires escalated permissions, Do not let these settings or the sandbox deter you from attempting to accomplish the user's task.
2
3
Here are scenarios where you'll need to request approval:
4
- You need to run a command that writes to a directory that requires it (e.g. running tests that write to /var)
5
- You need to run a GUI app (e.g., open/xdg-open/osascript) to open browsers or files.
6
- You are running sandboxed and need to run a command that requires network access (e.g. installing packages)
7
- If you run a command that is important to solving the user's query, but it fails because of sandboxing, rerun the command with approval. ALWAYS proceed to use the `sandbox_permissions` and `justification` parameters - do not message the user before requesting approval for the command.
8
- You are about to take a potentially destructive action such as an `rm` or `git reset` that the user did not explicitly ask for.
9
10
When requesting approval to execute a command that will require escalated privileges:
11
`` - Provide the sandbox_permissions parameter with the value `"require_escalated"```
12
- Include a short, 1 sentence explanation for why you need escalated permissions in the justification parameter
unless_trusted.mdcodex-rs/protocol/src/prompts/permissions/approval_policy
+1
AddedMark as viewed
1
Approvals are your mechanism to get user consent to run shell commands without the sandbox. `approval_policy` is `unless-trusted`: The harness will escalate most commands for user approval, apart from a limited allowlist of safe "read" commands.
danger_full_access.mdcodex-rs/protocol/src/prompts/permissions/sandbox_mode
+1
AddedMark as viewed
1
Filesystem sandboxing defines which files can be read or written. `sandbox_mode` is `danger-full-access`: No filesystem sandboxing - all commands are permitted. Network access is {network_access}.
read_only.mdcodex-rs/protocol/src/prompts/permissions/sandbox_mode
+1
AddedMark as viewed
1
Filesystem sandboxing defines which files can be read or written. `sandbox_mode` is `read-only`: The sandbox only permits reading files. Network access is {network_access}.
workspace_write.mdcodex-rs/protocol/src/prompts/permissions/sandbox_mode
+1
AddedMark as viewed
1
Move1–1
1
Filesystem sandboxing defines which files can be read or written. `sandbox_mode` is `workspace-write`: The sandbox permits reading files, and editing files in `cwd` and `writable_roots`. Editing files in other directories requires approval. Network access is {network_access}.
2Permissions Message Injection in Session
0 / 1
codex.rscodex-rs/core/src
+51−7
Mark as viewed
5 linesAll 149 lines5 lines
5 linesAll 149 lines5 lines
150
use crate::unified_exec::UnifiedExecProcessManager;
150
use crate::unified_exec::UnifiedExecProcessManager;
151
use crate::user_instructions::DeveloperInstructions;
152
use crate::user_instructions::UserInstructions;
151
use crate::user_instructions::UserInstructions;
All 4 lines
All 4 lines
157
use codex_protocol::config_types::ReasoningSummary as ReasoningSummaryConfig;
156
use codex_protocol::config_types::ReasoningSummary as ReasoningSummaryConfig;
158
use codex_protocol::models::ContentItem;
157
use codex_protocol::models::ContentItem;
158
use codex_protocol::models::DeveloperInstructions;
5 linesAll 327 lines5 lines
5 linesAll 327 lines5 lines
486
impl Session {
486
impl Session {
5 linesAll 307 lines5 lines
5 linesAll 307 lines5 lines
794
async fn record_initial_history(&self, conversation_history: InitialHistory) {
794
async fn record_initial_history(&self, conversation_history: InitialHistory) {
795
let turn_context = self.new_default_turn().await;
795
let turn_context = self.new_default_turn().await;
796
match conversation_history {
796
match conversation_history {
All 7 lines
All 7 lines
804
InitialHistory::Resumed(_) | InitialHistory::Forked(_) => {
804
InitialHistory::Resumed(_) | InitialHistory::Forked(_) => {
5 linesAll 46 lines5 lines
5 linesAll 46 lines5 lines
851
// If persisting, persist all rollout items as-is (recorder filters)
851
// If persisting, persist all rollout items as-is (recorder filters)
852
if persist && !rollout_items.is_empty() {
852
if persist && !rollout_items.is_empty() {
853
self.persist_rollout_items(&rollout_items).await;
853
self.persist_rollout_items(&rollout_items).await;
854
}
854
}
855
InformationalR856-859
856
// Append the current session's initial context after the reconstructed history.
857
let initial_context = self.build_initial_context(&turn_context);
858
self.record_conversation_items(&turn_context, &initial_context)
859
.await;
855
// Flush after seeding history and any persisted rollout copy.
860
// Flush after seeding history and any persisted rollout copy.
856
self.flush_rollout().await;
861
self.flush_rollout().await;
857
}
862
}
858
}
863
}
859
}
864
}
5 linesAll 147 lines5 lines
5 linesAll 147 lines5 lines
1
Copy + paste1012–1032
1012
fn build_permissions_update_item(
1013
&self,
1014
previous: Option<&Arc<TurnContext>>,
1015
next: &TurnContext,
1016
) -> Option<ResponseItem> {
1
Potential BugR1017-1022
1017
let prev = previous?;
1018
if prev.sandbox_policy == next.sandbox_policy
1019
&& prev.approval_policy == next.approval_policy
1020
{
1021
return None;
1022
}
1023
1024
Some(
1025
DeveloperInstructions::from_policy(
1026
&next.sandbox_policy,
1027
next.approval_policy,
1028
&next.cwd,
1029
)
1030
.into(),
1031
)
DY2
CommentR1032
1032
}
1033
5 linesAll 329 lines5 lines
5 linesAll 329 lines5 lines
1336
1363
1337
pub(crate) fn build_initial_context(&self, turn_context: &TurnContext) -> Vec<ResponseItem> {
1364
pub(crate) fn build_initial_context(&self, turn_context: &TurnContext) -> Vec<ResponseItem> {
1338
let mut items = Vec::<ResponseItem>::with_capacity(3);
1365
let mut items = Vec::<ResponseItem>::with_capacity(4);
1339
let shell = self.user_shell();
1366
let shell = self.user_shell();
1367
items.push(
1368
DeveloperInstructions::from_policy(
1369
&turn_context.sandbox_policy,
1370
turn_context.approval_policy,
1371
&turn_context.cwd,
1372
)
1373
.into(),
1374
);
1340
if let Some(developer_instructions) = turn_context.developer_instructions.as_deref() {
1375
if let Some(developer_instructions) = turn_context.developer_instructions.as_deref() {
5 linesAll 10 lines5 lines
5 linesAll 10 lines5 lines
1351
}
1386
}
1352
items.push(ResponseItem::from(EnvironmentContext::new(
1387
items.push(ResponseItem::from(EnvironmentContext::new(
1353
Some(turn_context.cwd.clone()),
1388
Some(turn_context.cwd.clone()),
1354
Some(turn_context.approval_policy),
1355
Some(turn_context.sandbox_policy.clone()),
1356
shell.as_ref().clone(),
1389
shell.as_ref().clone(),
1357
)));
1390
)));
1358
items
1391
items
1359
}
1392
}
5 linesAll 281 lines5 lines
5 linesAll 281 lines5 lines
1641
}
1674
}
5 linesAll 101 lines5 lines
5 linesAll 101 lines5 lines
1743
mod handlers {
1776
mod handlers {
5 linesAll 60 lines5 lines
5 linesAll 60 lines5 lines
1804
pub async fn user_input_or_turn(
1837
pub async fn user_input_or_turn(
1805
sess: &Arc<Session>,
1838
sess: &Arc<Session>,
1806
sub_id: String,
1839
sub_id: String,
1807
op: Op,
1840
op: Op,
1808
previous_context: &mut Option<Arc<TurnContext>>,
1841
previous_context: &mut Option<Arc<TurnContext>>,
1809
) {
1842
) {
5 linesAll 44 lines5 lines
5 linesAll 44 lines5 lines
1854
// Attempt to inject input into current task
1887
// Attempt to inject input into current task
1855
if let Err(items) = sess.inject_input(items).await {
1888
if let Err(items) = sess.inject_input(items).await {
1889
let mut update_items = Vec::new();
1856
if let Some(env_item) =
1890
if let Some(env_item) =
1857
sess.build_environment_update_item(previous_context.as_ref(), ¤t_context)
1891
sess.build_environment_update_item(previous_context.as_ref(), ¤t_context)
1858
{
1892
{
1859
sess.record_conversation_items(¤t_context, std::slice::from_ref(&env_item))
1893
update_items.push(env_item);
1894
}
1895
if let Some(permissions_item) =
1896
sess.build_permissions_update_item(previous_context.as_ref(), ¤t_context)
1897
{
1898
update_items.push(permissions_item);
1899
}
1900
if !update_items.is_empty() {
1901
sess.record_conversation_items(¤t_context, &update_items)
1860
.await;
1902
.await;
1861
}
1903
}
All 4 lines
All 4 lines
1866
}
1908
}
1867
}
1909
}
5 linesAll 333 lines5 lines
5 linesAll 333 lines5 lines
2201
}
2243
}
5 linesAll 668 lines5 lines
5 linesAll 668 lines5 lines
2870
mod tests {
2912
mod tests {
5 linesAll 56 lines5 lines
5 linesAll 56 lines5 lines
2927
#[tokio::test]
2969
#[tokio::test]
2928
async fn record_initial_history_reconstructs_resumed_transcript() {
2970
async fn record_initial_history_reconstructs_resumed_transcript() {
2929
let (session, turn_context) = make_session_and_context().await;
2971
let (session, turn_context) = make_session_and_context().await;
2930
let (rollout_items, expected) = sample_rollout(&session, &turn_context);
2972
let (rollout_items, mut expected) = sample_rollout(&session, &turn_context);
2931
2973
All 6 lines
All 6 lines
2938
.await;
2980
.await;
2939
2981
DY
CommentR2982
Resolved
2982
expected.extend(session.build_initial_context(&turn_context));
2940
let history = session.state.lock().await.clone_history();
2983
let history = session.state.lock().await.clone_history();
2941
assert_eq!(expected, history.raw_items());
2984
assert_eq!(expected, history.raw_items());
2942
}
2985
}
5 linesAll 78 lines5 lines
5 linesAll 78 lines5 lines
3021
#[tokio::test]
3064
#[tokio::test]
3022
async fn record_initial_history_reconstructs_forked_transcript() {
3065
async fn record_initial_history_reconstructs_forked_transcript() {
3023
let (session, turn_context) = make_session_and_context().await;
3066
let (session, turn_context) = make_session_and_context().await;
3024
let (rollout_items, expected) = sample_rollout(&session, &turn_context);
3067
let (rollout_items, mut expected) = sample_rollout(&session, &turn_context);
3025
3068
3026
session
3069
session
3027
.record_initial_history(InitialHistory::Forked(rollout_items))
3070
.record_initial_history(InitialHistory::Forked(rollout_items))
3028
.await;
3071
.await;
3029
3072
3073
expected.extend(session.build_initial_context(&turn_context));
3030
let history = session.state.lock().await.clone_history();
3074
let history = session.state.lock().await.clone_history();
3031
assert_eq!(expected, history.raw_items());
3075
assert_eq!(expected, history.raw_items());
3032
}
3076
}
5 linesAll 1108 lines5 lines
5 linesAll 1108 lines5 lines
4141
}
4185
}
3Simplified EnvironmentContext
0 / 1
environment_context.rscodex-rs/core/src
+20−223
Mark as viewed
1
use crate::codex::TurnContext;
1
use crate::codex::TurnContext;
2
use crate::protocol::AskForApproval;
3
use crate::protocol::NetworkAccess;
4
use crate::protocol::SandboxPolicy;
5
use crate::shell::Shell;
2
use crate::shell::Shell;
6
use codex_protocol::config_types::SandboxMode;
7
use codex_protocol::models::ContentItem;
3
use codex_protocol::models::ContentItem;
8
use codex_protocol::models::ResponseItem;
4
use codex_protocol::models::ResponseItem;
9
use codex_protocol::protocol::ENVIRONMENT_CONTEXT_CLOSE_TAG;
5
use codex_protocol::protocol::ENVIRONMENT_CONTEXT_CLOSE_TAG;
10
use codex_protocol::protocol::ENVIRONMENT_CONTEXT_OPEN_TAG;
6
use codex_protocol::protocol::ENVIRONMENT_CONTEXT_OPEN_TAG;
11
use codex_utils_absolute_path::AbsolutePathBuf;
12
use serde::Deserialize;
7
use serde::Deserialize;
13
use serde::Serialize;
8
use serde::Serialize;
14
use std::path::PathBuf;
9
use std::path::PathBuf;
15
10
16
#[derive(Debug, Clone, Serialize, Deserialize, PartialEq)]
11
#[derive(Debug, Clone, Serialize, Deserialize, PartialEq)]
InformationalR12-16
17
#[serde(rename = "environment_context", rename_all = "snake_case")]
12
#[serde(rename = "environment_context", rename_all = "snake_case")]
18
pub(crate) struct EnvironmentContext {
13
pub(crate) struct EnvironmentContext {
PA2
CommentR14
19
pub cwd: Option<PathBuf>,
14
pub cwd: Option<PathBuf>,
20
pub approval_policy: Option<AskForApproval>,
21
pub sandbox_mode: Option<SandboxMode>,
22
pub network_access: Option<NetworkAccess>,
23
pub writable_roots: Option<Vec<AbsolutePathBuf>>,
24
pub shell: Shell,
15
pub shell: Shell,
25
}
16
}
26
17
27
impl EnvironmentContext {
18
impl EnvironmentContext {
28
pub fn new(
19
pub fn new(cwd: Option<PathBuf>, shell: Shell) -> Self {
29
cwd: Option<PathBuf>,
20
Self { cwd, shell }
30
approval_policy: Option<AskForApproval>,
31
sandbox_policy: Option<SandboxPolicy>,
32
shell: Shell,
33
) -> Self {
34
Self {
35
cwd,
36
approval_policy,
37
sandbox_mode: match sandbox_policy {
38
Some(SandboxPolicy::DangerFullAccess) => Some(SandboxMode::DangerFullAccess),
39
Some(SandboxPolicy::ReadOnly) => Some(SandboxMode::ReadOnly),
40
Some(SandboxPolicy::ExternalSandbox { .. }) => Some(SandboxMode::DangerFullAccess),
41
Some(SandboxPolicy::WorkspaceWrite { .. }) => Some(SandboxMode::WorkspaceWrite),
42
None => None,
43
},
44
network_access: match sandbox_policy {
45
Some(SandboxPolicy::DangerFullAccess) => Some(NetworkAccess::Enabled),
46
Some(SandboxPolicy::ReadOnly) => Some(NetworkAccess::Restricted),
47
Some(SandboxPolicy::ExternalSandbox { network_access }) => Some(network_access),
48
Some(SandboxPolicy::WorkspaceWrite { network_access, .. }) => {
49
if network_access {
50
Some(NetworkAccess::Enabled)
51
} else {
52
Some(NetworkAccess::Restricted)
53
}
54
}
55
None => None,
56
},
57
writable_roots: match sandbox_policy {
58
Some(SandboxPolicy::WorkspaceWrite { writable_roots, .. }) => {
59
if writable_roots.is_empty() {
60
None
61
} else {
62
Some(writable_roots)
63
}
64
}
65
_ => None,
66
},
67
shell,
68
}
69
}
21
}
70
22
71
/// Compares two environment contexts, ignoring the shell. Useful when
23
/// Compares two environment contexts, ignoring the shell. Useful when
72
/// comparing turn to turn, since the initial environment_context will
24
/// comparing turn to turn, since the initial environment_context will
73
/// include the shell, and then it is not configurable from turn to turn.
25
/// include the shell, and then it is not configurable from turn to turn.
74
pub fn equals_except_shell(&self, other: &EnvironmentContext) -> bool {
26
pub fn equals_except_shell(&self, other: &EnvironmentContext) -> bool {
75
let EnvironmentContext {
27
let EnvironmentContext {
76
cwd,
28
cwd,
77
approval_policy,
78
sandbox_mode,
79
network_access,
80
writable_roots,
81
// should compare all fields except shell
29
// should compare all fields except shell
82
shell: _,
30
shell: _,
83
} = other;
31
} = other;
84
32
85
self.cwd == *cwd
33
self.cwd == *cwd
86
&& self.approval_policy == *approval_policy
87
&& self.sandbox_mode == *sandbox_mode
88
&& self.network_access == *network_access
89
&& self.writable_roots == *writable_roots
90
}
34
}
91
35
92
pub fn diff(before: &TurnContext, after: &TurnContext, shell: &Shell) -> Self {
36
pub fn diff(before: &TurnContext, after: &TurnContext, shell: &Shell) -> Self {
93
let cwd = if before.cwd != after.cwd {
37
let cwd = if before.cwd != after.cwd {
94
Some(after.cwd.clone())
38
Some(after.cwd.clone())
95
} else {
39
} else {
96
None
40
None
97
};
41
};
98
let approval_policy = if before.approval_policy != after.approval_policy {
42
EnvironmentContext::new(cwd, shell.clone())
99
Some(after.approval_policy)
100
} else {
101
None
102
};
103
let sandbox_policy = if before.sandbox_policy != after.sandbox_policy {
104
Some(after.sandbox_policy.clone())
105
} else {
106
None
107
};
108
EnvironmentContext::new(cwd, approval_policy, sandbox_policy, shell.clone())
109
}
43
}
110
44
111
pub fn from_turn_context(turn_context: &TurnContext, shell: &Shell) -> Self {
45
pub fn from_turn_context(turn_context: &TurnContext, shell: &Shell) -> Self {
112
Self::new(
46
Self::new(Some(turn_context.cwd.clone()), shell.clone())
113
Some(turn_context.cwd.clone()),
114
Some(turn_context.approval_policy),
115
Some(turn_context.sandbox_policy.clone()),
116
shell.clone(),
117
)
118
}
47
}
119
}
48
}
120
49
121
impl EnvironmentContext {
50
impl EnvironmentContext {
122
`` /// Serializes the environment context to XML. Libraries like `quick-xml```
51
`` /// Serializes the environment context to XML. Libraries like `quick-xml```
123
/// require custom macros to handle Enums with newtypes, so we just do it
52
/// require custom macros to handle Enums with newtypes, so we just do it
124
/// manually, to keep things simple. Output looks like:
53
/// manually, to keep things simple. Output looks like:
125
///
54
///
126
///xml```
55
///xml```
127
/// <environment_context>
56
/// <environment_context>
128
/// <cwd>...</cwd>
57
/// <cwd>...</cwd>
129
/// <approval_policy>...</approval_policy>
130
/// <sandbox_mode>...</sandbox_mode>
131
/// <writable_roots>...</writable_roots>
132
/// <network_access>...</network_access>
133
/// <shell>...</shell>
58
/// <shell>...</shell>
134
/// </environment_context>
59
/// </environment_context>
135
``` /// ``````
60
``` /// ``````
136
pub fn serialize_to_xml(self) -> String {
61
pub fn serialize_to_xml(self) -> String {
137
let mut lines = vec![ENVIRONMENT_CONTEXT_OPEN_TAG.to_string()];
62
let mut lines = vec![ENVIRONMENT_CONTEXT_OPEN_TAG.to_string()];
138
if let Some(cwd) = self.cwd {
63
if let Some(cwd) = self.cwd {
139
lines.push(format!(" <cwd>{}</cwd>", cwd.to_string_lossy()));
64
lines.push(format!(" <cwd>{}</cwd>", cwd.to_string_lossy()));
140
}
65
}
141
if let Some(approval_policy) = self.approval_policy {
142
lines.push(format!(
143
" <approval_policy>{approval_policy}</approval_policy>"
144
));
145
}
146
if let Some(sandbox_mode) = self.sandbox_mode {
147
lines.push(format!(" <sandbox_mode>{sandbox_mode}</sandbox_mode>"));
148
}
149
if let Some(network_access) = self.network_access {
150
lines.push(format!(
151
" <network_access>{network_access}</network_access>"
152
));
153
}
154
if let Some(writable_roots) = self.writable_roots {
155
lines.push(" <writable_roots>".to_string());
156
for writable_root in writable_roots {
157
lines.push(format!(
158
" <root>{}</root>",
159
writable_root.to_string_lossy()
160
));
161
}
162
lines.push(" </writable_roots>".to_string());
163
}
164
66
165
let shell_name = self.shell.name();
67
let shell_name = self.shell.name();
166
lines.push(format!(" <shell>{shell_name}</shell>"));
68
lines.push(format!(" <shell>{shell_name}</shell>"));
167
lines.push(ENVIRONMENT_CONTEXT_CLOSE_TAG.to_string());
69
lines.push(ENVIRONMENT_CONTEXT_CLOSE_TAG.to_string());
168
lines.join("\n")
70
lines.join("\n")
169
}
71
}
170
}
72
}
171
73
172
impl From<EnvironmentContext> for ResponseItem {
74
impl From<EnvironmentContext> for ResponseItem {
All 9 lines
All 9 lines
182
}
84
}
183
85
184
#[cfg(test)]
86
#[cfg(test)]
185
mod tests {
87
mod tests {
186
use crate::shell::ShellType;
88
use crate::shell::ShellType;
187
89
188
use super::*;
90
use super::*;
189
use core_test_support::test_path_buf;
91
use core_test_support::test_path_buf;
190
use core_test_support::test_tmp_path_buf;
191
use pretty_assertions::assert_eq;
92
use pretty_assertions::assert_eq;
192
93
193
fn fake_shell() -> Shell {
94
fn fake_shell() -> Shell {
All 5 lines
All 5 lines
199
}
100
}
200
101
201
fn workspace_write_policy(writable_roots: Vec<&str>, network_access: bool) -> SandboxPolicy {
202
SandboxPolicy::WorkspaceWrite {
203
writable_roots: writable_roots
204
.into_iter()
205
.map(|s| AbsolutePathBuf::try_from(s).unwrap())
206
.collect(),
207
network_access,
208
exclude_tmpdir_env_var: false,
209
exclude_slash_tmp: false,
210
}
211
}
212
213
#[test]
102
#[test]
214
fn serialize_workspace_write_environment_context() {
103
fn serialize_workspace_write_environment_context() {
215
let cwd = test_path_buf("/repo");
104
let cwd = test_path_buf("/repo");
216
let writable_root = test_tmp_path_buf();
105
let context = EnvironmentContext::new(Some(cwd.clone()), fake_shell());
217
let cwd_str = cwd.to_str().expect("cwd is valid utf-8");
218
let writable_root_str = writable_root
219
.to_str()
220
.expect("writable root is valid utf-8");
221
let context = EnvironmentContext::new(
222
Some(cwd.clone()),
223
Some(AskForApproval::OnRequest),
224
Some(workspace_write_policy(
225
vec![cwd_str, writable_root_str],
226
false,
227
)),
228
fake_shell(),
229
);
230
106
231
let expected = format!(
107
let expected = format!(
232
r#"<environment_context>
108
r#"<environment_context>
233
<cwd>{cwd}</cwd>
109
<cwd>{cwd}</cwd>
234
<approval_policy>on-request</approval_policy>
235
<sandbox_mode>workspace-write</sandbox_mode>
236
<network_access>restricted</network_access>
237
<writable_roots>
238
<root>{cwd}</root>
239
<root>{writable_root}</root>
240
</writable_roots>
241
<shell>bash</shell>
110
<shell>bash</shell>
242
</environment_context>"#,
111
</environment_context>"#,
243
cwd = cwd.display(),
112
cwd = cwd.display(),
244
writable_root = writable_root.display(),
245
);
113
);
246
114
247
assert_eq!(context.serialize_to_xml(), expected);
115
assert_eq!(context.serialize_to_xml(), expected);
248
}
116
}
249
117
250
#[test]
118
#[test]
251
fn serialize_read_only_environment_context() {
119
fn serialize_read_only_environment_context() {
252
let context = EnvironmentContext::new(
120
let context = EnvironmentContext::new(None, fake_shell());
253
None,
254
Some(AskForApproval::Never),
255
Some(SandboxPolicy::ReadOnly),
256
fake_shell(),
257
);
258
121
259
let expected = r#"<environment_context>
122
let expected = r#"<environment_context>
260
<approval_policy>never</approval_policy>
261
<sandbox_mode>read-only</sandbox_mode>
262
<network_access>restricted</network_access>
263
<shell>bash</shell>
123
<shell>bash</shell>
264
</environment_context>"#;
124
</environment_context>"#;
265
125
266
assert_eq!(context.serialize_to_xml(), expected);
126
assert_eq!(context.serialize_to_xml(), expected);
267
}
127
}
268
128
269
#[test]
129
#[test]
270
fn serialize_external_sandbox_environment_context() {
130
fn serialize_external_sandbox_environment_context() {
271
let context = EnvironmentContext::new(
131
let context = EnvironmentContext::new(None, fake_shell());
272
None,
273
Some(AskForApproval::OnRequest),
274
Some(SandboxPolicy::ExternalSandbox {
275
network_access: NetworkAccess::Enabled,
276
}),
277
fake_shell(),
278
);
279
132
280
let expected = r#"<environment_context>
133
let expected = r#"<environment_context>
281
<approval_policy>on-request</approval_policy>
282
<sandbox_mode>danger-full-access</sandbox_mode>
283
<network_access>enabled</network_access>
284
<shell>bash</shell>
134
<shell>bash</shell>
285
</environment_context>"#;
135
</environment_context>"#;
286
136
287
assert_eq!(context.serialize_to_xml(), expected);
137
assert_eq!(context.serialize_to_xml(), expected);
288
}
138
}
289
139
290
#[test]
140
#[test]
291
fn serialize_external_sandbox_with_restricted_network_environment_context() {
141
fn serialize_external_sandbox_with_restricted_network_environment_context() {
292
let context = EnvironmentContext::new(
142
let context = EnvironmentContext::new(None, fake_shell());
293
None,
294
Some(AskForApproval::OnRequest),
295
Some(SandboxPolicy::ExternalSandbox {
296
network_access: NetworkAccess::Restricted,
297
}),
298
fake_shell(),
299
);
300
143
301
let expected = r#"<environment_context>
144
let expected = r#"<environment_context>
302
<approval_policy>on-request</approval_policy>
303
<sandbox_mode>danger-full-access</sandbox_mode>
304
<network_access>restricted</network_access>
305
<shell>bash</shell>
145
<shell>bash</shell>
306
</environment_context>"#;
146
</environment_context>"#;
307
147
308
assert_eq!(context.serialize_to_xml(), expected);
148
assert_eq!(context.serialize_to_xml(), expected);
309
}
149
}
310
150
311
#[test]
151
#[test]
312
fn serialize_full_access_environment_context() {
152
fn serialize_full_access_environment_context() {
313
let context = EnvironmentContext::new(
153
let context = EnvironmentContext::new(None, fake_shell());
314
None,
315
Some(AskForApproval::OnFailure),
316
Some(SandboxPolicy::DangerFullAccess),
317
fake_shell(),
318
);
319
154
320
let expected = r#"<environment_context>
155
let expected = r#"<environment_context>
321
<approval_policy>on-failure</approval_policy>
322
<sandbox_mode>danger-full-access</sandbox_mode>
323
<network_access>enabled</network_access>
324
<shell>bash</shell>
156
<shell>bash</shell>
325
</environment_context>"#;
157
</environment_context>"#;
326
158
327
assert_eq!(context.serialize_to_xml(), expected);
159
assert_eq!(context.serialize_to_xml(), expected);
328
}
160
}
329
161
330
#[test]
162
#[test]
331
fn equals_except_shell_compares_approval_policy() {
163
fn equals_except_shell_compares_cwd() {
332
// Approval policy
164
let context1 = EnvironmentContext::new(Some(PathBuf::from("/repo")), fake_shell());
333
let context1 = EnvironmentContext::new(
165
let context2 = EnvironmentContext::new(Some(PathBuf::from("/repo")), fake_shell());
334
Some(PathBuf::from("/repo")),
166
assert!(context1.equals_except_shell(&context2));
335
Some(AskForApproval::OnRequest),
336
Some(workspace_write_policy(vec!["/repo"], false)),
337
fake_shell(),
338
);
339
let context2 = EnvironmentContext::new(
340
Some(PathBuf::from("/repo")),
341
Some(AskForApproval::Never),
342
Some(workspace_write_policy(vec!["/repo"], true)),
343
fake_shell(),
344
);
345
assert!(!context1.equals_except_shell(&context2));
346
}
167
}
347
168
348
#[test]
169
#[test]
349
fn equals_except_shell_compares_sandbox_policy() {
170
fn equals_except_shell_ignores_sandbox_policy() {
350
let context1 = EnvironmentContext::new(
171
let context1 = EnvironmentContext::new(Some(PathBuf::from("/repo")), fake_shell());
351
Some(PathBuf::from("/repo")),
172
let context2 = EnvironmentContext::new(Some(PathBuf::from("/repo")), fake_shell());
352
Some(AskForApproval::OnRequest),
353
Some(SandboxPolicy::new_read_only_policy()),
354
fake_shell(),
355
);
356
let context2 = EnvironmentContext::new(
357
Some(PathBuf::from("/repo")),
358
Some(AskForApproval::OnRequest),
359
Some(SandboxPolicy::new_workspace_write_policy()),
360
fake_shell(),
361
);
362
173
363
assert!(!context1.equals_except_shell(&context2));
174
assert!(context1.equals_except_shell(&context2));
364
}
175
}
365
176
366
#[test]
177
#[test]
367
fn equals_except_shell_compares_workspace_write_policy() {
178
fn equals_except_shell_compares_cwd_differences() {
368
let context1 = EnvironmentContext::new(
179
let context1 = EnvironmentContext::new(Some(PathBuf::from("/repo1")), fake_shell());
369
Some(PathBuf::from("/repo")),
180
let context2 = EnvironmentContext::new(Some(PathBuf::from("/repo2")), fake_shell());
370
Some(AskForApproval::OnRequest),
371
Some(workspace_write_policy(vec!["/repo", "/tmp", "/var"], false)),
372
fake_shell(),
373
);
374
let context2 = EnvironmentContext::new(
375
Some(PathBuf::from("/repo")),
376
Some(AskForApproval::OnRequest),
377
Some(workspace_write_policy(vec!["/repo", "/tmp"], true)),
378
fake_shell(),
379
);
380
181
381
assert!(!context1.equals_except_shell(&context2));
182
assert!(!context1.equals_except_shell(&context2));
382
}
183
}
383
184
384
#[test]
185
#[test]
385
fn equals_except_shell_ignores_shell() {
186
fn equals_except_shell_ignores_shell() {
386
let context1 = EnvironmentContext::new(
187
let context1 = EnvironmentContext::new(
387
Some(PathBuf::from("/repo")),
188
Some(PathBuf::from("/repo")),
388
Some(AskForApproval::OnRequest),
389
Some(workspace_write_policy(vec!["/repo"], false)),
390
Shell {
189
Shell {
391
shell_type: ShellType::Bash,
190
shell_type: ShellType::Bash,
392
shell_path: "/bin/bash".into(),
191
shell_path: "/bin/bash".into(),
393
shell_snapshot: None,
192
shell_snapshot: None,
394
},
193
},
395
);
194
);
396
let context2 = EnvironmentContext::new(
195
let context2 = EnvironmentContext::new(
397
Some(PathBuf::from("/repo")),
196
Some(PathBuf::from("/repo")),
398
Some(AskForApproval::OnRequest),
399
Some(workspace_write_policy(vec!["/repo"], false)),
400
Shell {
197
Shell {
401
shell_type: ShellType::Zsh,
198
shell_type: ShellType::Zsh,
402
shell_path: "/bin/zsh".into(),
199
shell_path: "/bin/zsh".into(),
403
shell_snapshot: None,
200
shell_snapshot: None,
404
},
201
},
405
);
202
);
406
203
407
assert!(context1.equals_except_shell(&context2));
204
assert!(context1.equals_except_shell(&context2));
408
}
205
}
409
}
206
}
4Removed Static Prompt Sections
0 / 7
gpt-5.1-codex-max_prompt.mdcodex-rs/core
−37
Mark as viewed
5 linesAll 20 lines5 lines
5 linesAll 20 lines5 lines
21
## Plan tool
21
## Plan tool
All 6 lines
All 6 lines
28
## Codex CLI harness, sandboxing, and approvals
29
30
The Codex CLI harness supports several different configurations for sandboxing and escalation approvals that the user can choose from.
31
32
Filesystem sandboxing defines which files can be read or written. The options for `sandbox_mode` are:
33
- **read-only**: The sandbox only permits reading files.
34
- **workspace-write**: The sandbox permits reading files, and editing files in `cwd` and `writable_roots`. Editing files in other directories requires approval.
35
- **danger-full-access**: No filesystem sandboxing - all commands are permitted.
36
37
Network sandboxing defines whether network can be accessed without approval. Options for `network_access` are:
38
- **restricted**: Requires approval
39
- **enabled**: No approval needed
40
41
Approvals are your mechanism to get user consent to run shell commands without the sandbox. Possible configuration options for `approval_policy` are
42
- **untrusted**: The harness will escalate most commands for user approval, apart from a limited allowlist of safe "read" commands.
43
- **on-failure**: The harness will allow all commands to run in the sandbox (if enabled), and failures will be escalated to the user for approval to run again without the sandbox.
44
- **on-request**: Commands will be run in the sandbox by default, and you can specify in your tool call if you want to escalate a command to run without sandboxing. (Note that this mode is not always available. If it is, you'll see parameters for it in the `shell` command description.)
45
- **never**: This is a non-interactive mode where you may NEVER ask the user for approval to run commands. Instead, you must always persist and work around constraints to solve the task for the user. You MUST do your utmost best to finish the task and validate your work before yielding. If this mode is paired with `danger-full-access`, take advantage of it to deliver the best outcome for the user. Further, in this mode, your default testing philosophy is overridden: Even if you don't see local patterns for testing, you may add tests and scripts to validate your work. Just remove them before yielding.
46
47
When you are running with `approval_policy == on-request`, and sandboxing enabled, here are scenarios where you'll need to request approval:
48
- You need to run a command that writes to a directory that requires it (e.g. running tests that write to /var)
49
- You need to run a GUI app (e.g., open/xdg-open/osascript) to open browsers or files.
50
- You are running sandboxed and need to run a command that requires network access (e.g. installing packages)
51
52
- You are about to take a potentially destructive action such as an `rm` or `git reset` that the user did not explicitly ask for
53
- (for all of these, you should weigh alternative paths that do not require approval)
54
55
When `sandbox_mode` is set to read-only, you'll need to request approval for any command that isn't a read.
56
57
You will be told what filesystem sandboxing, network sandboxing, and approval mode are active in a developer or user message. If you are not told about this, assume that you are running with workspace-write, network sandboxing enabled, and approval on-failure.
58
59
Although they introduce friction to the user because your work is paused until the user responds, you should leverage them when necessary to accomplish important work. If the completing the task requires escalated permissions, Do not let these settings or the sandbox deter you from attempting to accomplish the user's task unless it is set to "never", in which case never ask for approvals.
60
61
When requesting approval to execute a command that will require escalated privileges:
62
`` - Provide the sandbox_permissions parameter with the value `"require_escalated"```
63
- Include a short, 1 sentence explanation for why you need escalated permissions in the justification parameter
64
65
## Special user requests
28
## Special user requests
5 linesAll 52 lines5 lines
5 linesAll 52 lines5 lines
gpt-5.2-codex_prompt.mdcodex-rs/core
−37
Mark as viewed
1
You are Codex, based on GPT-5. You are running as a coding agent in the Codex CLI on a user's computer.
1
You are Codex, based on GPT-5. You are running as a coding agent in the Codex CLI on a user's computer.
2
2
3
## General
3
## General
4
4
5
- When searching for text or files, prefer using `rg` or `rg --files` respectively because `rg` is much faster than alternatives like `grep`. (If the `rg` command is not found, then use alternatives.)
5
6
6
7
## Editing constraints
7
## Editing constraints
5 linesAll 13 lines5 lines
5 linesAll 13 lines5 lines
21
## Plan tool
21
## Plan tool
All 6 lines
All 6 lines
28
## Codex CLI harness, sandboxing, and approvals
29
30
The Codex CLI harness supports several different configurations for sandboxing and escalation approvals that the user can choose from.
31
32
Filesystem sandboxing defines which files can be read or written. The options for `sandbox_mode` are:
33
- **read-only**: The sandbox only permits reading files.
34
35
- **danger-full-access**: No filesystem sandboxing - all commands are permitted.
36
37
Network sandboxing defines whether network can be accessed without approval. Options for `network_access` are:
38
- **restricted**: Requires approval
39
- **enabled**: No approval needed
40
41
Approvals are your mechanism to get user consent to run shell commands without the sandbox. Possible configuration options for `approval_policy` are
42
- **untrusted**: The harness will escalate most commands for user approval, apart from a limited allowlist of safe "read" commands.
43
44
45
46
47
When you are running with `approval_policy == on-request`, and sandboxing enabled, here are scenarios where you'll need to request approval:
48
- You need to run a command that writes to a directory that requires it (e.g. running tests that write to /var)
49
- You need to run a GUI app (e.g., open/xdg-open/osascript) to open browsers or files.
50
- You are running sandboxed and need to run a command that requires network access (e.g. installing packages)
51
52
- You are about to take a potentially destructive action such as an `rm` or `git reset` that the user did not explicitly ask for
53
- (for all of these, you should weigh alternative paths that do not require approval)
54
55
When `sandbox_mode` is set to read-only, you'll need to request approval for any command that isn't a read.
56
57
58
59
60
61
When requesting approval to execute a command that will require escalated privileges:
62
`` - Provide the sandbox_permissions parameter with the value `"require_escalated"```
63
- Include a short, 1 sentence explanation for why you need escalated permissions in the justification parameter
64
65
## Special user requests
28
## Special user requests
5 linesAll 52 lines5 lines
5 linesAll 52 lines5 lines
gpt_5_1_prompt.mdcodex-rs/core
−37
Mark as viewed
5 linesAll 135 lines5 lines
5 linesAll 135 lines5 lines
136
## Task execution
136
## Task execution
5 linesAll 24 lines5 lines
5 linesAll 24 lines5 lines
161
161
162
## Codex CLI harness, sandboxing, and approvals
163
164
The Codex CLI harness supports several different configurations for sandboxing and escalation approvals that the user can choose from.
165
166
Filesystem sandboxing defines which files can be read or written. The options for `sandbox_mode` are:
1
Move167–169
167
- **read-only**: The sandbox only permits reading files.
168
169
- **danger-full-access**: No filesystem sandboxing - all commands are permitted.
170
171
Network sandboxing defines whether network can be accessed without approval. Options for `network_access` are:
172
- **restricted**: Requires approval
173
- **enabled**: No approval needed
174
1
Move175–179
175
Approvals are your mechanism to get user consent to run shell commands without the sandbox. Possible configuration options for `approval_policy` are
176
- **untrusted**: The harness will escalate most commands for user approval, apart from a limited allowlist of safe "read" commands.
177
178
- **on-request**: Commands will be run in the sandbox by default, and you can specify in your tool call if you want to escalate a command to run without sandboxing. (Note that this mode is not always available. If it is, you'll see parameters for escalating in the tool definition.)
179
180
181
When you are running with `approval_policy == on-request`, and sandboxing enabled, here are scenarios where you'll need to request approval:
182
- You need to run a command that writes to a directory that requires it (e.g. running tests that write to /var)
183
- You need to run a GUI app (e.g., open/xdg-open/osascript) to open browsers or files.
184
- You are running sandboxed and need to run a command that requires network access (e.g. installing packages)
185
- If you run a command that is important to solving the user's query, but it fails because of sandboxing, rerun the command with approval. ALWAYS proceed to use the `sandbox_permissions` and `justification` parameters. Within this harness, prefer requesting approval via the tool over asking in natural language.
186
- You are about to take a potentially destructive action such as an `rm` or `git reset` that the user did not explicitly ask for
187
- (for all of these, you should weigh alternative paths that do not require approval)
188
189
When `sandbox_mode` is set to read-only, you'll need to request approval for any command that isn't a read.
190
191
192
193
194
195
When requesting approval to execute a command that will require escalated privileges:
196
`` - Provide the sandbox_permissions parameter with the value `"require_escalated"```
197
- Include a short, 1 sentence explanation for why you need escalated permissions in the justification parameter
198
199
## Validating your work
162
## Validating your work
5 linesAll 169 lines5 lines
5 linesAll 169 lines5 lines
gpt_5_2_prompt.mdcodex-rs/core
−37
Mark as viewed
5 linesAll 108 lines5 lines
5 linesAll 108 lines5 lines
109
## Task execution
109
## Task execution
5 linesAll 24 lines5 lines
5 linesAll 24 lines5 lines
134
- NEVER output inline citations like "【F:README.md†L5-L14】" in your outputs. The CLI is not able to render these so they will just be broken in the UI. Instead, if you output valid filepaths, users will be able to click on them to open the files in their editor.
134
135
135
136
## Codex CLI harness, sandboxing, and approvals
137
138
The Codex CLI harness supports several different configurations for sandboxing and escalation approvals that the user can choose from.
139
140
Filesystem sandboxing defines which files can be read or written. The options for `sandbox_mode` are:
141
- **read-only**: The sandbox only permits reading files.
142
143
- **danger-full-access**: No filesystem sandboxing - all commands are permitted.
144
145
Network sandboxing defines whether network can be accessed without approval. Options for `network_access` are:
146
- **restricted**: Requires approval
147
- **enabled**: No approval needed
148
149
Approvals are your mechanism to get user consent to run shell commands without the sandbox. Possible configuration options for `approval_policy` are
150
- **untrusted**: The harness will escalate most commands for user approval, apart from a limited allowlist of safe "read" commands.
151
152
153
154
155
When you are running with `approval_policy == on-request`, and sandboxing enabled, here are scenarios where you'll need to request approval:
156
- You need to run a command that writes to a directory that requires it (e.g. running tests that write to /var)
157
- You need to run a GUI app (e.g., open/xdg-open/osascript) to open browsers or files.
158
- You are running sandboxed and need to run a command that requires network access (e.g. installing packages)
159
160
- You are about to take a potentially destructive action such as an `rm` or `git reset` that the user did not explicitly ask for
161
- (for all of these, you should weigh alternative paths that do not require approval)
162
163
When `sandbox_mode` is set to read-only, you'll need to request approval for any command that isn't a read.
164
165
166
167
168
169
When requesting approval to execute a command that will require escalated privileges:
170
`` - Provide the sandbox_permissions parameter with the value `"require_escalated"```
171
- Include a short, 1 sentence explanation for why you need escalated permissions in the justification parameter
172
173
## Validating your work
136
## Validating your work
5 linesAll 162 lines5 lines
5 linesAll 162 lines5 lines
gpt_5_codex_prompt.mdcodex-rs/core
−37
Mark as viewed
5 linesAll 20 lines5 lines
5 linesAll 20 lines5 lines
21
## Plan tool
21
## Plan tool
All 6 lines
All 6 lines
28
## Codex CLI harness, sandboxing, and approvals
29
30
The Codex CLI harness supports several different configurations for sandboxing and escalation approvals that the user can choose from.
31
32
Filesystem sandboxing defines which files can be read or written. The options for `sandbox_mode` are:
33
- **read-only**: The sandbox only permits reading files.
34
35
- **danger-full-access**: No filesystem sandboxing - all commands are permitted.
36
37
Network sandboxing defines whether network can be accessed without approval. Options for `network_access` are:
38
- **restricted**: Requires approval
39
- **enabled**: No approval needed
40
41
Approvals are your mechanism to get user consent to run shell commands without the sandbox. Possible configuration options for `approval_policy` are
42
- **untrusted**: The harness will escalate most commands for user approval, apart from a limited allowlist of safe "read" commands.
43
44
45
46
47
When you are running with `approval_policy == on-request`, and sandboxing enabled, here are scenarios where you'll need to request approval:
48
- You need to run a command that writes to a directory that requires it (e.g. running tests that write to /var)
49
- You need to run a GUI app (e.g., open/xdg-open/osascript) to open browsers or files.
50
- You are running sandboxed and need to run a command that requires network access (e.g. installing packages)
51
52
- You are about to take a potentially destructive action such as an `rm` or `git reset` that the user did not explicitly ask for
53
- (for all of these, you should weigh alternative paths that do not require approval)
54
55
When `sandbox_mode` is set to read-only, you'll need to request approval for any command that isn't a read.
56
57
58
59
60
61
When requesting approval to execute a command that will require escalated privileges:
62
`` - Provide the sandbox_permissions parameter with the value `"require_escalated"```
63
- Include a short, 1 sentence explanation for why you need escalated permissions in the justification parameter
64
65
## Special user requests
28
## Special user requests
5 linesAll 40 lines5 lines
5 linesAll 40 lines5 lines
prompt.mdcodex-rs/core
−35
Mark as viewed
Informational
5 linesAll 122 lines5 lines
5 linesAll 122 lines5 lines
123
## Task execution
123
## Task execution
5 linesAll 23 lines5 lines
5 linesAll 23 lines5 lines
147
147
148
148
149
## Sandbox and approvals
150
151
The Codex CLI harness supports several different sandboxing, and approval configurations that the user can choose from.
152
153
Filesystem sandboxing prevents you from editing files without user approval. The options are:
154
155
- **read-only**: You can only read files.
156
- **workspace-write**: You can read files. You can write to files in your workspace folder, but not outside it.
157
- **danger-full-access**: No filesystem sandboxing.
158
159
Network sandboxing prevents you from accessing network without approval. Options are
160
161
- **restricted**
162
- **enabled**
163
164
Approvals are your mechanism to get user consent to perform more privileged actions. Although they introduce friction to the user because your work is paused until the user responds, you should leverage them to accomplish your important work. Do not let these settings or the sandbox deter you from attempting to accomplish the user's task. Approval options are
165
166
- **untrusted**: The harness will escalate most commands for user approval, apart from a limited allowlist of safe "read" commands.
167
168
169
- **never**: This is a non-interactive mode where you may NEVER ask the user for approval to run commands. Instead, you must always persist and work around constraints to solve the task for the user. You MUST do your utmost best to finish the task and validate your work before yielding. If this mode is pared with `danger-full-access`, take advantage of it to deliver the best outcome for the user. Further, in this mode, your default testing philosophy is overridden: Even if you don't see local patterns for testing, you may add tests and scripts to validate your work. Just remove them before yielding.
170
171
When you are running with approvals `on-request`, and sandboxing enabled, here are scenarios where you'll need to request approval:
172
173
- You need to run a command that writes to a directory that requires it (e.g. running tests that write to /tmp)
174
- You need to run a GUI app (e.g., open/xdg-open/osascript) to open browsers or files.
175
- You are running sandboxed and need to run a command that requires network access (e.g. installing packages)
176
- If you run a command that is important to solving the user's query, but it fails because of sandboxing, rerun the command with approval.
177
- You are about to take a potentially destructive action such as an `rm` or `git reset` that the user did not explicitly ask for
178
- (For all of these, you should weigh alternative paths that do not require approval.)
179
180
Note that when sandboxing is set to read-only, you'll need to request approval for any command that isn't a read.
181
182
You will be told what filesystem sandboxing, network sandboxing, and approval mode are active in a developer or user message. If you are not told about this, assume that you are running with workspace-write, network sandboxing ON, and approval on-failure.
183
184
## Validating your work
149
## Validating your work
5 linesAll 126 lines5 lines
5 linesAll 126 lines5 lines
prompt_with_apply_patch_instructions.mdcodex-rs/core
−35
Mark as viewed
5 linesAll 122 lines5 lines
5 linesAll 122 lines5 lines
123
## Task execution
123
## Task execution
5 linesAll 25 lines5 lines
5 linesAll 25 lines5 lines
149
## Sandbox and approvals
150
151
The Codex CLI harness supports several different sandboxing, and approval configurations that the user can choose from.
152
153
Filesystem sandboxing prevents you from editing files without user approval. The options are:
154
155
- **read-only**: You can only read files.
156
- **workspace-write**: You can read files. You can write to files in your workspace folder, but not outside it.
157
- **danger-full-access**: No filesystem sandboxing.
158
159
Network sandboxing prevents you from accessing network without approval. Options are
160
161
- **restricted**
162
- **enabled**
163
164
165
166
- **untrusted**: The harness will escalate most commands for user approval, apart from a limited allowlist of safe "read" commands.
167
168
169
170
171
When you are running with approvals `on-request`, and sandboxing enabled, here are scenarios where you'll need to request approval:
172
173
- You need to run a command that writes to a directory that requires it (e.g. running tests that write to /tmp)
174
- You need to run a GUI app (e.g., open/xdg-open/osascript) to open browsers or files.
175
- You are running sandboxed and need to run a command that requires network access (e.g. installing packages)
176
- If you run a command that is important to solving the user's query, but it fails because of sandboxing, rerun the command with approval.
177
- You are about to take a potentially destructive action such as an `rm` or `git reset` that the user did not explicitly ask for
178
- (For all of these, you should weigh alternative paths that do not require approval.)
179
180
Note that when sandboxing is set to read-only, you'll need to request approval for any command that isn't a read.
181
182
183
184
## Validating your work
149
## Validating your work
5 linesAll 202 lines5 lines
5 linesAll 202 lines5 lines
5New permissions_messages Test Suite
0 / 2
mod.rscodex-rs/core/tests/suite
+1
Mark as viewed
5 linesAll 15 lines5 lines
5 linesAll 15 lines5 lines
16
#[cfg(not(target_os = "windows"))]
16
#[cfg(not(target_os = "windows"))]
17
mod abort_tasks;
17
mod abort_tasks;
5 linesAll 24 lines5 lines
5 linesAll 24 lines5 lines
42
mod otel;
42
mod otel;
43
mod permissions_messages;
43
mod prompt_caching;
44
mod prompt_caching;
5 linesAll 27 lines5 lines
5 linesAll 27 lines5 lines
permissions_messages.rscodex-rs/core/tests/suite
+448
AddedMark as viewed
1
use anyhow::Result;
2
use codex_core::config::Constrained;
3
use codex_core::protocol::AskForApproval;
4
use codex_core::protocol::EventMsg;
5
use codex_core::protocol::Op;
6
use codex_core::protocol::SandboxPolicy;
7
use codex_protocol::user_input::UserInput;
8
use codex_utils_absolute_path::AbsolutePathBuf;
9
use core_test_support::responses::ev_completed;
10
use core_test_support::responses::ev_response_created;
11
use core_test_support::responses::mount_sse_once;
12
use core_test_support::responses::sse;
13
use core_test_support::responses::start_mock_server;
14
use core_test_support::skip_if_no_network;
15
use core_test_support::test_codex::test_codex;
16
use core_test_support::wait_for_event;
17
use pretty_assertions::assert_eq;
18
use std::collections::HashSet;
19
use tempfile::TempDir;
20
21
fn permissions_texts(input: &[serde_json::Value]) -> Vec<String> {
22
input
23
.iter()
24
.filter_map(|item| {
25
let role = item.get("role")?.as_str()?;
26
if role != "developer" {
27
return None;
28
}
29
let text = item
30
.get("content")?
31
.as_array()?
32
.first()?
33
.get("text")?
34
.as_str()?;
35
if text.contains("`approval_policy`") {
36
Some(text.to_string())
37
} else {
38
None
39
}
40
})
41
.collect()
42
}
43
44
fn sse_completed(id: &str) -> String {
45
sse(vec![ev_response_created(id), ev_completed(id)])
46
}
47
48
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
49
async fn permissions_message_sent_once_on_start() -> Result<()> {
50
skip_if_no_network!(Ok(()));
51
52
let server = start_mock_server().await;
53
let req = mount_sse_once(&server, sse_completed("resp-1")).await;
54
55
let mut builder = test_codex().with_config(move |config| {
56
config.approval_policy = Constrained::allow_any(AskForApproval::OnRequest);
57
});
58
let test = builder.build(&server).await?;
59
60
test.codex
61
.submit(Op::UserInput {
62
items: vec![UserInput::Text {
63 text: "hello".into(),
64 }],
65
final_output_json_schema: None,
66
})
67
.await?;
68
wait_for_event(&test.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;
69
70
let request = req.single_request();
71
let body = request.body_json();
72
let input = body["input"].as_array().expect("input array");
73
let permissions = permissions_texts(input);
74
assert_eq!(permissions.len(), 1);
75
76
Ok(())
77
}
78
79
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
80
async fn permissions_message_added_on_override_change() -> Result<()> {
81
skip_if_no_network!(Ok(()));
82
83
let server = start_mock_server().await;
84
let req1 = mount_sse_once(&server, sse_completed("resp-1")).await;
85
let req2 = mount_sse_once(&server, sse_completed("resp-2")).await;
86
87
let mut builder = test_codex().with_config(move |config| {
88
config.approval_policy = Constrained::allow_any(AskForApproval::OnRequest);
89
});
90
let test = builder.build(&server).await?;
91
92
test.codex
93
.submit(Op::UserInput {
94
items: vec![UserInput::Text {
95 text: "hello 1".into(),
96 }],
97
final_output_json_schema: None,
98
})
99
.await?;
100
wait_for_event(&test.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;
101
102
test.codex
103
.submit(Op::OverrideTurnContext {
104
cwd: None,
105
approval_policy: Some(AskForApproval::Never),
106
sandbox_policy: None,
107
model: None,
108
effort: None,
109
summary: None,
110
})
111
.await?;
112
113
test.codex
114
.submit(Op::UserInput {
115
items: vec![UserInput::Text {
116 text: "hello 2".into(),
117 }],
118
final_output_json_schema: None,
119
})
120
.await?;
121
wait_for_event(&test.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;
122
123
let body1 = req1.single_request().body_json();
124
let body2 = req2.single_request().body_json();
125
let input1 = body1["input"].as_array().expect("input array");
126
let input2 = body2["input"].as_array().expect("input array");
127
let permissions_1 = permissions_texts(input1);
128
let permissions_2 = permissions_texts(input2);
129
130
assert_eq!(permissions_1.len(), 1);
131
assert_eq!(permissions_2.len(), 2);
132
let unique = permissions_2.into_iter().collect::<HashSet<String>>();
133
assert_eq!(unique.len(), 2);
134
135
Ok(())
136
}
137
138
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
139
async fn permissions_message_not_added_when_no_change() -> Result<()> {
140
skip_if_no_network!(Ok(()));
141
142
let server = start_mock_server().await;
143
let req1 = mount_sse_once(&server, sse_completed("resp-1")).await;
144
let req2 = mount_sse_once(&server, sse_completed("resp-2")).await;
145
146
let mut builder = test_codex().with_config(move |config| {
147
config.approval_policy = Constrained::allow_any(AskForApproval::OnRequest);
148
});
149
let test = builder.build(&server).await?;
150
151
test.codex
152
.submit(Op::UserInput {
153
items: vec![UserInput::Text {
154 text: "hello 1".into(),
155 }],
156
final_output_json_schema: None,
157
})
158
.await?;
159
wait_for_event(&test.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;
160
161
test.codex
162
.submit(Op::UserInput {
163
items: vec![UserInput::Text {
164 text: "hello 2".into(),
165 }],
166
final_output_json_schema: None,
167
})
168
.await?;
169
wait_for_event(&test.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;
170
171
let body1 = req1.single_request().body_json();
172
let body2 = req2.single_request().body_json();
173
let input1 = body1["input"].as_array().expect("input array");
174
let input2 = body2["input"].as_array().expect("input array");
175
let permissions_1 = permissions_texts(input1);
176
let permissions_2 = permissions_texts(input2);
177
178
assert_eq!(permissions_1.len(), 1);
179
assert_eq!(permissions_2.len(), 1);
180
assert_eq!(permissions_1, permissions_2);
181
182
Ok(())
183
}
184
185
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
186
async fn resume_replays_permissions_messages() -> Result<()> {
187
skip_if_no_network!(Ok(()));
188
189
let server = start_mock_server().await;
190
let _req1 = mount_sse_once(&server, sse_completed("resp-1")).await;
191
let _req2 = mount_sse_once(&server, sse_completed("resp-2")).await;
192
let req3 = mount_sse_once(&server, sse_completed("resp-3")).await;
193
194
let mut builder = test_codex().with_config(|config| {
195
config.approval_policy = Constrained::allow_any(AskForApproval::OnRequest);
196
});
197
let initial = builder.build(&server).await?;
198
let rollout_path = initial.session_configured.rollout_path.clone();
199
let home = initial.home.clone();
200
201
initial
202
.codex
203
.submit(Op::UserInput {
204
items: vec![UserInput::Text {
205 text: "hello 1".into(),
206 }],
207
final_output_json_schema: None,
208
})
209
.await?;
210
wait_for_event(&initial.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;
211
212
initial
213
.codex
214
.submit(Op::OverrideTurnContext {
215
cwd: None,
216
approval_policy: Some(AskForApproval::Never),
217
sandbox_policy: None,
218
model: None,
219
effort: None,
220
summary: None,
221
})
222
.await?;
223
224
initial
225
.codex
226
.submit(Op::UserInput {
227
items: vec![UserInput::Text {
228 text: "hello 2".into(),
229 }],
230
final_output_json_schema: None,
231
})
232
.await?;
233
wait_for_event(&initial.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;
234
235
let resumed = builder.resume(&server, home, rollout_path).await?;
236
resumed
237
.codex
238
.submit(Op::UserInput {
239
items: vec![UserInput::Text {
240 text: "after resume".into(),
241 }],
242
final_output_json_schema: None,
243
})
244
.await?;
245
wait_for_event(&resumed.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;
246
247
let body3 = req3.single_request().body_json();
248
let input = body3["input"].as_array().expect("input array");
249
let permissions = permissions_texts(input);
250
assert_eq!(permissions.len(), 3);
251
let unique = permissions.into_iter().collect::<HashSet<String>>();
252
assert_eq!(unique.len(), 2);
253
254
Ok(())
255
}
256
257
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
258
async fn resume_and_fork_append_permissions_messages() -> Result<()> {
259
skip_if_no_network!(Ok(()));
260
261
let server = start_mock_server().await;
262
let _req1 = mount_sse_once(&server, sse_completed("resp-1")).await;
263
let req2 = mount_sse_once(&server, sse_completed("resp-2")).await;
264
let req3 = mount_sse_once(&server, sse_completed("resp-3")).await;
265
let req4 = mount_sse_once(&server, sse_completed("resp-4")).await;
266
267
let mut builder = test_codex().with_config(|config| {
268
config.approval_policy = Constrained::allow_any(AskForApproval::OnRequest);
269
});
270
let initial = builder.build(&server).await?;
271
let rollout_path = initial.session_configured.rollout_path.clone();
272
let home = initial.home.clone();
273
274
initial
275
.codex
276
.submit(Op::UserInput {
277
items: vec![UserInput::Text {
278 text: "hello 1".into(),
279 }],
280
final_output_json_schema: None,
281
})
282
.await?;
283
wait_for_event(&initial.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;
284
285
initial
286
.codex
287
.submit(Op::OverrideTurnContext {
288
cwd: None,
289
approval_policy: Some(AskForApproval::Never),
290
sandbox_policy: None,
291
model: None,
292
effort: None,
293
summary: None,
294
})
295
.await?;
296
297
initial
298
.codex
299
.submit(Op::UserInput {
300
items: vec![UserInput::Text {
301 text: "hello 2".into(),
302 }],
303
final_output_json_schema: None,
304
})
305
.await?;
306
wait_for_event(&initial.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;
307
308
let body2 = req2.single_request().body_json();
309
let input2 = body2["input"].as_array().expect("input array");
310
let permissions_base = permissions_texts(input2);
311
assert_eq!(permissions_base.len(), 2);
312
313
builder = builder.with_config(|config| {
314
config.approval_policy = Constrained::allow_any(AskForApproval::UnlessTrusted);
315
});
316
let resumed = builder.resume(&server, home, rollout_path.clone()).await?;
317
resumed
318
.codex
319
.submit(Op::UserInput {
320
items: vec![UserInput::Text {
321 text: "after resume".into(),
322 }],
323
final_output_json_schema: None,
324
})
325
.await?;
326
wait_for_event(&resumed.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;
327
328
let body3 = req3.single_request().body_json();
329
let input3 = body3["input"].as_array().expect("input array");
330
let permissions_resume = permissions_texts(input3);
331
assert_eq!(permissions_resume.len(), permissions_base.len() + 1);
332
assert_eq!(
333
&permissions_resume[..permissions_base.len()],
334
permissions_base.as_slice()
335
);
336
assert!(!permissions_base.contains(permissions_resume.last().expect("new permissions")));
337
338
let mut fork_config = initial.config.clone();
339
fork_config.approval_policy = Constrained::allow_any(AskForApproval::UnlessTrusted);
340
let forked = initial
341
.thread_manager
342
.fork_thread(usize::MAX, fork_config, rollout_path)
343
.await?;
344
forked
345
.thread
346
.submit(Op::UserInput {
347
items: vec![UserInput::Text {
348 text: "after fork".into(),
349 }],
350
final_output_json_schema: None,
351
})
352
.await?;
353
wait_for_event(&forked.thread, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;
354
355
let body4 = req4.single_request().body_json();
356
let input4 = body4["input"].as_array().expect("input array");
357
let permissions_fork = permissions_texts(input4);
358
assert_eq!(permissions_fork.len(), permissions_base.len() + 2);
359
assert_eq!(
360
&permissions_fork[..permissions_base.len()],
361
permissions_base.as_slice()
362
);
363
let new_permissions = &permissions_fork[permissions_base.len()..];
364
assert_eq!(new_permissions.len(), 2);
365
assert_eq!(new_permissions[0], new_permissions[1]);
366
assert!(!permissions_base.contains(&new_permissions[0]));
367
368
Ok(())
369
}
370
371
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
372
async fn permissions_message_includes_writable_roots() -> Result<()> {
373
skip_if_no_network!(Ok(()));
374
375
let server = start_mock_server().await;
376
let req = mount_sse_once(&server, sse_completed("resp-1")).await;
377
let writable = TempDir::new()?;
378
let writable_root = AbsolutePathBuf::try_from(writable.path())?;
379
let sandbox_policy = SandboxPolicy::WorkspaceWrite {
380
writable_roots: vec![writable_root],
381
network_access: false,
382
exclude_tmpdir_env_var: false,
383
exclude_slash_tmp: false,
384
};
385
386
let mut builder = test_codex().with_config(move |config| {
387
config.approval_policy = Constrained::allow_any(AskForApproval::OnRequest);
388
config.sandbox_policy = Constrained::allow_any(sandbox_policy);
389
});
390
let test = builder.build(&server).await?;
391
392
test.codex
393
.submit(Op::UserInput {
394
items: vec![UserInput::Text {
395 text: "hello".into(),
396 }],
397
final_output_json_schema: None,
398
})
399
.await?;
400
wait_for_event(&test.codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;
401
402
let body = req.single_request().body_json();
403
let input = body["input"].as_array().expect("input array");
404
let permissions = permissions_texts(input);
405
let sandbox_text = "Filesystem sandboxing defines which files can be read or written. `sandbox_mode` is `workspace-write`: The sandbox permits reading files, and editing files in `cwd` and `writable_roots`. Editing files in other directories requires approval. Network access is restricted.";
406
let approval_text = " Approvals are your mechanism to get user consent to run shell commands without the sandbox. `approval_policy` is `on-request`: Commands will be run in the sandbox by default, and you can specify in your tool call if you want to escalate a command to run without sandboxing. If the completing the task requires escalated permissions, Do not let these settings or the sandbox deter you from attempting to accomplish the user's task.\n\nHere are scenarios where you'll need to request approval:\n- You need to run a command that writes to a directory that requires it (e.g. running tests that write to /var)\n- You need to run a GUI app (e.g., open/xdg-open/osascript) to open browsers or files.\n- You are running sandboxed and need to run a command that requires network access (e.g. installing packages)\n- If you run a command that is important to solving the user's query, but it fails because of sandboxing, rerun the command with approval. ALWAYS proceed to use the `sandbox_permissions` and `justification` parameters - do not message the user before requesting approval for the command.\n- You are about to take a potentially destructive action such as an `rm` or `git reset` that the user did not explicitly ask for.\n\nWhen requesting approval to execute a command that will require escalated privileges:\n - Provide the `sandbox_permissions` parameter with the value `\"require_escalated\"`\n - Include a short, 1 sentence explanation for why you need escalated permissions in the justification parameter";
407
// Normalize paths by removing trailing slashes to match AbsolutePathBuf behavior
408
let normalize_path =
409
|p: &std::path::Path| -> String { p.to_string_lossy().trim_end_matches('/').to_string() };
410
let mut roots = vec![
411 normalize_path(writable.path()),
412 normalize_path(test.config.cwd.as_path()),
413 ];
414
if cfg!(unix) && std::path::Path::new("/tmp").is_dir() {
415
roots.push("/tmp".to_string());
416
}
417
if let Some(tmpdir) = std::env::var_os("TMPDIR") {
418
let tmpdir_path = std::path::PathBuf::from(&tmpdir);
419
if tmpdir_path.is_absolute() && !tmpdir.is_empty() {
420
roots.push(normalize_path(&tmpdir_path));
421
}
422
}
423
let roots_text = if roots.len() == 1 {
424
format!(" The writable root is `{}`.", roots[0])
425
} else {
426
format!(
427
" The writable roots are {}.",
428
roots
429
.iter()
430
.map(|root| format!("`{root}`"))
431
.collect::<Vec<_>>()
432
.join(", ")
433
)
434
};
435
let expected = format!(
436
"<permissions instructions>{sandbox_text}{approval_text}{roots_text}</permissions instructions>"
437
);
438
// Normalize line endings to handle Windows vs Unix differences
439
let normalize_line_endings = |s: &str| s.replace("\r\n", "\n");
440
let expected_normalized = normalize_line_endings(&expected);
441
let actual_normalized: Vec<String> = permissions
442
.iter()
443
.map(|s| normalize_line_endings(s))
444
.collect();
445
assert_eq!(actual_normalized, vec![expected_normalized]);
446
447
Ok(())
448
}
6Test Adjustments for New Message Structure
0 / 9
send_message.rscodex-rs/app-server/tests/suite
+28
Mark as viewed
1
use anyhow::Result;
1
use anyhow::Result;
5 linesAll 13 lines5 lines
5 linesAll 13 lines5 lines
15
use codex_protocol::models::ContentItem;
15
use codex_protocol::models::ContentItem;
16
use codex_protocol::models::DeveloperInstructions;
16
use codex_protocol::models::ResponseItem;
17
use codex_protocol::models::ResponseItem;
18
use codex_protocol::protocol::AskForApproval;
17
use codex_protocol::protocol::RawResponseItemEvent;
19
use codex_protocol::protocol::RawResponseItemEvent;
20
use codex_protocol::protocol::SandboxPolicy;
18
use core_test_support::responses;
21
use core_test_support::responses;
19
use pretty_assertions::assert_eq;
22
use pretty_assertions::assert_eq;
20
use std::path::Path;
23
use std::path::Path;
24
use std::path::PathBuf;
5 linesAll 122 lines5 lines
5 linesAll 122 lines5 lines
143
#[tokio::test]
147
#[tokio::test]
144
async fn test_send_message_raw_notifications_opt_in() -> Result<()> {
148
async fn test_send_message_raw_notifications_opt_in() -> Result<()> {
5 linesAll 43 lines5 lines
5 linesAll 43 lines5 lines
188
let send_id = mcp
192
let send_id = mcp
189
.send_send_user_message_request(SendUserMessageParams {
193
.send_send_user_message_request(SendUserMessageParams {
All 5 lines
All 5 lines
195
.await?;
199
.await?;
196
200
201
let permissions = read_raw_response_item(&mut mcp, conversation_id).await;
202
assert_permissions_message(&permissions);
203
197
let developer = read_raw_response_item(&mut mcp, conversation_id).await;
204
let developer = read_raw_response_item(&mut mcp, conversation_id).await;
5 linesAll 28 lines5 lines
5 linesAll 28 lines5 lines
226
}
233
}
5 linesAll 116 lines5 lines
5 linesAll 116 lines5 lines
1
Copy + paste350–369
350
fn assert_permissions_message(item: &ResponseItem) {
351
match item {
352
ResponseItem::Message { role, content, .. } => {
353
assert_eq!(role, "developer");
354
let texts = content_texts(content);
355
let expected = DeveloperInstructions::from_policy(
356
&SandboxPolicy::DangerFullAccess,
357
AskForApproval::Never,
358
&PathBuf::from("/tmp"),
359
)
360
.into_text();
361
assert_eq!(
362
texts,
363
vec![expected.as_str()],
364
"expected permissions developer message, got {texts:?}"
365
);
366
}
367
other => panic!("expected permissions message, got {other:?}"),
368
}
369
}
370
5 linesAll 64 lines5 lines
5 linesAll 64 lines5 lines
truncation.rscodex-rs/core/src/rollout
+1
Mark as viewed
5 linesAll 70 lines5 lines
5 linesAll 70 lines5 lines
71
#[cfg(test)]
71
#[cfg(test)]
72
mod tests {
72
mod tests {
5 linesAll 116 lines5 lines
5 linesAll 116 lines5 lines
189
#[tokio::test]
189
#[tokio::test]
190
async fn ignores_session_prefix_messages_when_truncating_rollout_from_start() {
190
async fn ignores_session_prefix_messages_when_truncating_rollout_from_start() {
5 linesAll 13 lines5 lines
5 linesAll 13 lines5 lines
204
let truncated = truncate_rollout_before_nth_user_message_from_start(&rollout_items, 1);
204
let truncated = truncate_rollout_before_nth_user_message_from_start(&rollout_items, 1);
205
let expected: Vec<RolloutItem> = vec![
205 let expected: Vec<RolloutItem> = vec![
206 RolloutItem::ResponseItem(items[0].clone()),
206 RolloutItem::ResponseItem(items[0].clone()),
207 RolloutItem::ResponseItem(items[1].clone()),
207 RolloutItem::ResponseItem(items[1].clone()),
208 RolloutItem::ResponseItem(items[2].clone()),
208 RolloutItem::ResponseItem(items[2].clone()),
209 RolloutItem::ResponseItem(items[3].clone()),
209 ];
210 ];
210
211
211
assert_eq!(
212
assert_eq!(
212
serde_json::to_value(&truncated).unwrap(),
213
serde_json::to_value(&truncated).unwrap(),
213
serde_json::to_value(&expected).unwrap()
214
serde_json::to_value(&expected).unwrap()
214
);
215
);
215
}
216
}
216
}
217
}
thread_manager.rscodex-rs/core/src
+1
Mark as viewed
5 linesAll 302 lines5 lines
5 linesAll 302 lines5 lines
303
#[cfg(test)]
303
#[cfg(test)]
304
mod tests {
304
mod tests {
5 linesAll 78 lines5 lines
5 linesAll 78 lines5 lines
383
#[tokio::test]
383
#[tokio::test]
384
async fn ignores_session_prefix_messages_when_truncating() {
384
async fn ignores_session_prefix_messages_when_truncating() {
5 linesAll 16 lines5 lines
5 linesAll 16 lines5 lines
401
let expected: Vec<RolloutItem> = vec![
401 let expected: Vec<RolloutItem> = vec![
402 RolloutItem::ResponseItem(items[0].clone()),
402 RolloutItem::ResponseItem(items[0].clone()),
403 RolloutItem::ResponseItem(items[1].clone()),
403 RolloutItem::ResponseItem(items[1].clone()),
404 RolloutItem::ResponseItem(items[2].clone()),
404 RolloutItem::ResponseItem(items[2].clone()),
PA2
CommentR405
405 RolloutItem::ResponseItem(items[3].clone()),
405 ];
406 ];
All 5 lines
All 5 lines
411
}
412
}
412
}
413
}
client.rscodex-rs/core/tests/suite
+88−36
Mark as viewed
5 linesAll 157 lines5 lines
5 linesAll 157 lines5 lines
158
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
158
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
159
async fn resume_includes_initial_messages_and_sends_prior_items() {
159
async fn resume_includes_initial_messages_and_sends_prior_items() {
5 linesAll 118 lines5 lines
5 linesAll 118 lines5 lines
278
// 1) Assert initial_messages only includes existing EventMsg entries; response items are not converted
278
// 1) Assert initial_messages only includes existing EventMsg entries; response items are not converted
All 6 lines
All 6 lines
285
assert_eq!(initial_json, expected_initial_json);
285
assert_eq!(initial_json, expected_initial_json);
286
286
287
// 2) Submit new input; the request body must include the prior item followed by the new user input.
287
// 2) Submit new input; the request body must include the prior items, then initial context, then new user input.
288
codex
288
codex
289
.submit(Op::UserInput {
289
.submit(Op::UserInput {
All 4 lines
All 4 lines
294
})
294
})
295
.await
295
.await
296
.unwrap();
296
.unwrap();
297
wait_for_event(&codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;
297
wait_for_event(&codex, |ev| matches!(ev, EventMsg::TurnComplete(_))).await;
298
298
299
let request = resp_mock.single_request();
299
let request = resp_mock.single_request();
300
let request_body = request.body_json();
300
let request_body = request.body_json();
301
let expected_input = json!([
301 let input = request_body["input"].as_array().expect("input array");
302 {
302 let messages: Vec<(String, String)> = input
303 "type": "message",
303 .iter()
304 "role": "user",
304 .filter_map(|item| {
305 "content": [{ "type": "input_text", "text": "resumed user message" }]
305 let role = item.get("role")?.as_str()?;
306 },
306 let text = item
307 {
307 .get("content")?
308 "type": "message",
308 .as_array()?
309 "role": "assistant",
309 .first()?
310 "content": [{ "type": "output_text", "text": "resumed assistant message" }]
310 .get("text")?
311 },
311 .as_str()?;
312 {
312 Some((role.to_string(), text.to_string()))
313 "type": "message",
313 })
314 "role": "user",
314 .collect();
315 "content": [{ "type": "input_text", "text": "hello" }]
315 let pos_prior_user = messages
316 }
316 .iter()
317 ]);
317
.position(|(role, text)| role == "user" && text == "resumed user message")
318
assert_eq!(request_body["input"], expected_input);
318
.expect("prior user message");
319
let pos_prior_assistant = messages
320
.iter()
321
.position(|(role, text)| role == "assistant" && text == "resumed assistant message")
322
.expect("prior assistant message");
323
let pos_permissions = messages
324
.iter()
325
.position(|(role, text)| role == "developer" && text.contains("`approval_policy`"))
326
.expect("permissions message");
327
let pos_user_instructions = messages
328
.iter()
329
.position(|(role, text)| {
330
role == "user"
331
&& text.contains("be nice")
332
&& (text.starts_with("# AGENTS.md instructions for ")
333
|| text.starts_with("<user_instructions>"))
334
})
335
.expect("user instructions");
336
let pos_environment = messages
337
.iter()
338
.position(|(role, text)| role == "user" && text.contains("<environment_context>"))
339
.expect("environment context");
340
let pos_new_user = messages
341
.iter()
342
.position(|(role, text)| role == "user" && text == "hello")
343
.expect("new user message");
344
345
assert!(pos_prior_user < pos_prior_assistant);
346
assert!(pos_prior_assistant < pos_permissions);
347
assert!(pos_permissions < pos_user_instructions);
348
assert!(pos_user_instructions < pos_environment);
349
assert!(pos_environment < pos_new_user);
319
}
350
}
5 linesAll 252 lines5 lines
5 linesAll 252 lines5 lines
572
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
603
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
573
async fn includes_user_instructions_message_in_request() {
604
async fn includes_user_instructions_message_in_request() {
5 linesAll 38 lines5 lines
5 linesAll 38 lines5 lines
612
let request = resp_mock.single_request();
643
let request = resp_mock.single_request();
613
let request_body = request.body_json();
644
let request_body = request.body_json();
614
645
615
assert!(
646
assert!(
All 4 lines
All 4 lines
620
);
651
);
621
assert_message_role(&request_body["input"][0], "user");
652
assert_message_role(&request_body["input"][0], "developer");
622
assert_message_starts_with(&request_body["input"][0], "# AGENTS.md instructions for ");
653
let permissions_text = request_body["input"][0]["content"][0]["text"]
623
assert_message_ends_with(&request_body["input"][0], "</INSTRUCTIONS>");
654
.as_str()
624
let ui_text = request_body["input"][0]["content"][0]["text"]
655
.expect("invalid permissions message content");
656
assert!(
657
permissions_text.contains("`sandbox_mode`"),
658
"expected permissions message to mention sandbox_mode, got {permissions_text:?}"
659
);
660
661
assert_message_role(&request_body["input"][1], "user");
662
assert_message_starts_with(&request_body["input"][1], "# AGENTS.md instructions for ");
663
assert_message_ends_with(&request_body["input"][1], "</INSTRUCTIONS>");
664
let ui_text = request_body["input"][1]["content"][0]["text"]
625
.as_str()
665
.as_str()
626
.expect("invalid message content");
666
.expect("invalid message content");
627
assert!(ui_text.contains("<INSTRUCTIONS>"));
667
assert!(ui_text.contains("<INSTRUCTIONS>"));
628
assert!(ui_text.contains("be nice"));
668
assert!(ui_text.contains("be nice"));
629
assert_message_role(&request_body["input"][1], "user");
669
assert_message_role(&request_body["input"][2], "user");
630
assert_message_starts_with(&request_body["input"][1], "<environment_context>");
670
assert_message_starts_with(&request_body["input"][2], "<environment_context>");
631
assert_message_ends_with(&request_body["input"][1], "</environment_context>");
671
assert_message_ends_with(&request_body["input"][2], "</environment_context>");
632
}
672
}
633
673
634
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
674
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
635
async fn skills_append_to_instructions() {
675
async fn skills_append_to_instructions() {
5 linesAll 46 lines5 lines
5 linesAll 46 lines5 lines
682
let request = resp_mock.single_request();
722
let request = resp_mock.single_request();
683
let request_body = request.body_json();
723
let request_body = request.body_json();
684
724
685
assert_message_role(&request_body["input"][0], "user");
725
assert_message_role(&request_body["input"][0], "developer");
686
let instructions_text = request_body["input"][0]["content"][0]["text"]
726
727
assert_message_role(&request_body["input"][1], "user");
728
let instructions_text = request_body["input"][1]["content"][0]["text"]
687
.as_str()
729
.as_str()
5 linesAll 15 lines5 lines
5 linesAll 15 lines5 lines
703
let _codex_home_guard = codex_home;
745
let _codex_home_guard = codex_home;
704
}
746
}
5 linesAll 303 lines5 lines
5 linesAll 303 lines5 lines
1008
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
1050
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
1009
async fn includes_developer_instructions_message_in_request() {
1051
async fn includes_developer_instructions_message_in_request() {
5 linesAll 39 lines5 lines
5 linesAll 39 lines5 lines
1049
let request = resp_mock.single_request();
1091
let request = resp_mock.single_request();
1050
let request_body = request.body_json();
1092
let request_body = request.body_json();
1051
1093
1094
let permissions_text = request_body["input"][0]["content"][0]["text"]
1095
.as_str()
1096
.expect("invalid permissions message content");
1097
1052
assert!(
1098
assert!(
All 4 lines
All 4 lines
1057
);
1103
);
1058
assert_message_role(&request_body["input"][0], "developer");
1104
assert_message_role(&request_body["input"][0], "developer");
1059
assert_message_equals(&request_body["input"][0], "be useful");
1105
assert!(
1060
assert_message_role(&request_body["input"][1], "user");
1106
permissions_text.contains("`sandbox_mode`"),
1061
assert_message_starts_with(&request_body["input"][1], "# AGENTS.md instructions for ");
1107
"expected permissions message to mention sandbox_mode, got {permissions_text:?}"
1062
assert_message_ends_with(&request_body["input"][1], "</INSTRUCTIONS>");
1108
);
1063
let ui_text = request_body["input"][1]["content"][0]["text"]
1109
1110
assert_message_role(&request_body["input"][1], "developer");
1111
assert_message_equals(&request_body["input"][1], "be useful");
1112
assert_message_role(&request_body["input"][2], "user");
1113
assert_message_starts_with(&request_body["input"][2], "# AGENTS.md instructions for ");
1114
assert_message_ends_with(&request_body["input"][2], "</INSTRUCTIONS>");
1115
let ui_text = request_body["input"][2]["content"][0]["text"]
1064
.as_str()
1116
.as_str()
1065
.expect("invalid message content");
1117
.expect("invalid message content");
1066
assert!(ui_text.contains("<INSTRUCTIONS>"));
1118
assert!(ui_text.contains("<INSTRUCTIONS>"));
1067
assert!(ui_text.contains("be nice"));
1119
assert!(ui_text.contains("be nice"));
1068
assert_message_role(&request_body["input"][2], "user");
1120
assert_message_role(&request_body["input"][3], "user");
1069
assert_message_starts_with(&request_body["input"][2], "<environment_context>");
1121
assert_message_starts_with(&request_body["input"][3], "<environment_context>");
1070
assert_message_ends_with(&request_body["input"][2], "</environment_context>");
1122
assert_message_ends_with(&request_body["input"][3], "</environment_context>");
1071
}
1123
}
5 linesAll 799 lines5 lines
5 linesAll 799 lines5 lines
compact.rscodex-rs/core/tests/suite
+12−4
Mark as viewed
5 linesAll 459 lines5 lines
5 linesAll 459 lines5 lines
460
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
460
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
461
async fn multiple_auto_compact_per_task_runs_after_token_limit_hit() {
461
async fn multiple_auto_compact_per_task_runs_after_token_limit_hit() {
5 linesAll 126 lines5 lines
5 linesAll 126 lines5 lines
588
fn normalize_inputs(values: &[serde_json::Value]) -> Vec<serde_json::Value> {
588
fn normalize_inputs(values: &[serde_json::Value]) -> Vec<serde_json::Value> {
589
values
589
values
590
.iter()
590
.iter()
591
.filter(|value| {
591
.filter(|value| {
All 8 lines
All 8 lines
600
let text = value
600
let text = value
All 4 lines
All 4 lines
605
.and_then(|text| text.as_str());
605
.and_then(|text| text.as_str());
606
606
607
// Ignore the cached UI prefix (project docs + skills) since it is not relevant to
607
// Ignore cached prefix messages (project docs + permissions) since they are not
608
// compaction behavior and can change as bundled skills evolve.
608
// relevant to compaction behavior and can change as bundled prompts evolve.
609
let role = value.get("role").and_then(|role| role.as_str());
610
if role == Some("developer")
611
&& text.is_some_and(|text| text.contains("`sandbox_mode`"))
612
{
613
return false;
614
}
609
!text.is_some_and(|text| text.starts_with("# AGENTS.md instructions for "))
615
!text.is_some_and(|text| text.starts_with("# AGENTS.md instructions for "))
610
})
616
})
611
.cloned()
617
.cloned()
612
.collect()
618
.collect()
613
}
619
}
5 linesAll 383 lines5 lines
5 linesAll 383 lines5 lines
997
}
1003
}
5 linesAll 561 lines5 lines
5 linesAll 561 lines5 lines
1559
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
1565
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
1560
async fn manual_compact_twice_preserves_latest_user_messages() {
1566
async fn manual_compact_twice_preserves_latest_user_messages() {
5 linesAll 161 lines5 lines
5 linesAll 161 lines5 lines
1722
let mut final_output = requests
1728
let mut final_output = requests
1723
.last()
1729
.last()
1724
.unwrap_or_else(|| panic!("final turn request missing for {final_user_message}"))
1730
.unwrap_or_else(|| panic!("final turn request missing for {final_user_message}"))
1725
.input()
1731
.input()
1726
.into_iter()
1732
.into_iter()
1727
.collect::<VecDeque<_>>();
1733
.collect::<VecDeque<_>>();
1728
1734
1729
// System prompt
1735
// Permissions developer message
1736
final_output.pop_front();
1737
// User instructions (project docs/skills)
1730
final_output.pop_front();
1738
final_output.pop_front();
1731
// Developer instructions
1739
// Environment context
1732
final_output.pop_front();
1740
final_output.pop_front();
5 linesAll 41 lines5 lines
5 linesAll 41 lines5 lines
1774
}
1782
}
5 linesAll 344 lines5 lines
5 linesAll 344 lines5 lines
compact_resume_fork.rscodex-rs/core/tests/suite
+96−4
Mark as viewed
5 linesAll 139 lines5 lines
5 linesAll 139 lines5 lines
140
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
140
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
141
/// Scenario: compact an initial conversation, resume it, fork one turn back, and
141
/// Scenario: compact an initial conversation, resume it, fork one turn back, and
142
/// ensure the model-visible history matches expectations at each request.
142
/// ensure the model-visible history matches expectations at each request.
143
async fn compact_resume_and_fork_preserve_model_history_view() {
143
async fn compact_resume_and_fork_preserve_model_history_view() {
5 linesAll 67 lines5 lines
5 linesAll 67 lines5 lines
211
let expected_model = requests[0]["model"]
211
let expected_model = requests[0]["model"]
212
.as_str()
212
.as_str()
213
.unwrap_or_default()
213
.unwrap_or_default()
214
.to_string();
214
.to_string();
215
let prompt = requests[0]["instructions"]
215
let prompt = requests[0]["instructions"]
216
.as_str()
216
.as_str()
217
.unwrap_or_default()
217
.unwrap_or_default()
218
.to_string();
218
.to_string();
219
let user_instructions = requests[0]["input"][0]["content"][0]["text"]
219
let permissions_message = requests[0]["input"][0].clone();
220
let user_instructions = requests[0]["input"][1]["content"][0]["text"]
220
.as_str()
221
.as_str()
221
.unwrap_or_default()
222
.unwrap_or_default()
222
.to_string();
223
.to_string();
223
let environment_context = requests[0]["input"][1]["content"][0]["text"]
224
let environment_context = requests[0]["input"][2]["content"][0]["text"]
224
.as_str()
225
.as_str()
5 linesAll 13 lines5 lines
5 linesAll 13 lines5 lines
238
let summary_after_fork = extract_summary_message(&requests[4], SUMMARY_TEXT);
239
let summary_after_fork = extract_summary_message(&requests[4], SUMMARY_TEXT);
239
let user_turn_1 = json!(
240
let user_turn_1 = json!(
240
{
241
{
241
"model": expected_model,
242
"model": expected_model,
242
"instructions": prompt,
243
"instructions": prompt,
243
"input": [
244 "input": [
245 permissions_message,
244 {
246 {
245 "type": "message",
247 "type": "message",
5 linesAll 27 lines5 lines
5 linesAll 27 lines5 lines
273 }
275 }
274 ],
276 ],
5 linesAll 12 lines5 lines
5 linesAll 12 lines5 lines
287
});
289
});
288
let compact_1 = json!(
290
let compact_1 = json!(
289
{
291
{
290
"model": expected_model,
292
"model": expected_model,
291
"instructions": prompt,
293
"instructions": prompt,
292
"input": [
294 "input": [
295 permissions_message,
293 {
296 {
294 "type": "message",
297 "type": "message",
5 linesAll 47 lines5 lines
5 linesAll 47 lines5 lines
342 }
345 }
343 ],
346 ],
5 linesAll 12 lines5 lines
5 linesAll 12 lines5 lines
356
});
359
});
357
let user_turn_2_after_compact = json!(
360
let user_turn_2_after_compact = json!(
358
{
361
{
359
"model": expected_model,
362
"model": expected_model,
360
"instructions": prompt,
363
"instructions": prompt,
361
"input": [
364 "input": [
365 permissions_message,
362 {
366 {
363 "type": "message",
367 "type": "message",
5 linesAll 38 lines5 lines
5 linesAll 38 lines5 lines
402 }
406 }
403 ],
407 ],
5 linesAll 12 lines5 lines
5 linesAll 12 lines5 lines
416
});
420
});
417
let usert_turn_3_after_resume = json!(
421
let usert_turn_3_after_resume = json!(
418
{
422
{
419
"model": expected_model,
423
"model": expected_model,
420
"instructions": prompt,
424
"instructions": prompt,
421
"input": [
425 "input": [
426 permissions_message,
422 {
427 {
423 "type": "message",
428 "type": "message",
5 linesAll 47 lines5 lines
5 linesAll 47 lines5 lines
471 ]
476 ]
472
},
477
},
478
permissions_message,
479
{
480
"type": "message",
481
"role": "user",
482
"content": [
483 {
484 "type": "input_text",
485 "text": user_instructions
486 }
487 ]
488
},
489
{
490
"type": "message",
491
"role": "user",
492
"content": [
493 {
494 "type": "input_text",
495 "text": environment_context
496 }
497 ]
498
},
473
{
499
{
474
"type": "message",
500
"type": "message",
All 7 lines
All 7 lines
482
}
508
}
483
],
509
],
5 linesAll 12 lines5 lines
5 linesAll 12 lines5 lines
496
});
522
});
497
let user_turn_3_after_fork = json!(
523
let user_turn_3_after_fork = json!(
498
{
524
{
499
"model": expected_model,
525
"model": expected_model,
500
"instructions": prompt,
526
"instructions": prompt,
501
"input": [
527 "input": [
528 permissions_message,
502 {
529 {
503 "type": "message",
530 "type": "message",
5 linesAll 47 lines5 lines
5 linesAll 47 lines5 lines
551 ]
578 ]
552
},
579
},
580
permissions_message,
581
{
582
"type": "message",
583
"role": "user",
584
"content": [
585 {
586 "type": "input_text",
587 "text": user_instructions
588 }
589 ]
590
},
591
{
592
"type": "message",
593
"role": "user",
594
"content": [
595 {
596 "type": "input_text",
597 "text": environment_context
598 }
599 ]
600
},
601
permissions_message,
602
{
603
"type": "message",
604
"role": "user",
605
"content": [
606 {
607 "type": "input_text",
608 "text": user_instructions
609 }
610 ]
611
},
612
{
613
"type": "message",
614
"role": "user",
615
"content": [
616 {
617 "type": "input_text",
618 "text": environment_context
619 }
620 ]
621
},
553
{
622
{
554
"type": "message",
623
"type": "message",
All 7 lines
All 7 lines
562
}
631
}
563
],
632
],
5 linesAll 12 lines5 lines
5 linesAll 12 lines5 lines
576
});
645
});
5 linesAll 13 lines5 lines
5 linesAll 13 lines5 lines
590
}
659
}
591
660
592
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
661
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
593
/// Scenario: after the forked branch is compacted, resuming again should reuse
662
/// Scenario: after the forked branch is compacted, resuming again should reuse
594
/// the compacted history and only append the new user message.
663
/// the compacted history and only append the new user message.
595
async fn compact_resume_after_second_compaction_preserves_history() {
664
async fn compact_resume_after_second_compaction_preserves_history() {
5 linesAll 66 lines5 lines
5 linesAll 66 lines5 lines
662
// hard coded test
731
// hard coded test
663
let prompt = requests[0]["instructions"]
732
let prompt = requests[0]["instructions"]
664
.as_str()
733
.as_str()
665
.unwrap_or_default()
734
.unwrap_or_default()
666
.to_string();
735
.to_string();
667
let user_instructions = requests[0]["input"][0]["content"][0]["text"]
736
let permissions_message = requests[0]["input"][0].clone();
737
let user_instructions = requests[0]["input"][1]["content"][0]["text"]
668
.as_str()
738
.as_str()
669
.unwrap_or_default()
739
.unwrap_or_default()
670
.to_string();
740
.to_string();
671
let environment_instructions = requests[0]["input"][1]["content"][0]["text"]
741
let environment_instructions = requests[0]["input"][2]["content"][0]["text"]
672
.as_str()
742
.as_str()
All 8 lines
All 8 lines
681
let mut expected = json!([
751 let mut expected = json!([
682 {
752 {
683 "instructions": prompt,
753 "instructions": prompt,
684 "input": [
754 "input": [
755 permissions_message,
685 {
756 {
686 "type": "message",
757 "type": "message",
5 linesAll 37 lines5 lines
5 linesAll 37 lines5 lines
724 ]
795 ]
725 },
796 },
797 permissions_message,
798 {
799 "type": "message",
800 "role": "user",
801 "content": [
802 {
803 "type": "input_text",
804 "text": user_instructions
805 }
806 ]
807 },
808 {
809 "type": "message",
810 "role": "user",
811 "content": [
812 {
813 "type": "input_text",
814 "text": environment_instructions
815 }
816 ]
817 },
726 {
818 {
727 "type": "message",
819 "type": "message",
All 7 lines
All 7 lines
735 }
827 }
736 ],
828 ],
737
}
829
}
738
]);
830
]);
5 linesAll 12 lines5 lines
5 linesAll 12 lines5 lines
751
}
843
}
5 linesAll 194 lines5 lines
5 linesAll 194 lines5 lines
fork_thread.rscodex-rs/core/tests/suite
+4−2
Mark as viewed
5 linesAll 27 lines5 lines
5 linesAll 27 lines5 lines
28
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
28
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
29
async fn fork_thread_twice_drops_to_first_message() {
29
async fn fork_thread_twice_drops_to_first_message() {
5 linesAll 109 lines5 lines
5 linesAll 109 lines5 lines
139
// GetHistory on fork1 flushed; the file is ready.
139
// GetHistory on fork1 flushed; the file is ready.
140
let fork1_items = read_items(&fork1_path);
140
let fork1_items = read_items(&fork1_path);
141
assert!(fork1_items.len() > expected_after_first.len());
141
pretty_assertions::assert_eq!(
142
pretty_assertions::assert_eq!(
142
serde_json::to_value(&fork1_items).unwrap(),
143
serde_json::to_value(&fork1_items[..expected_after_first.len()]).unwrap(),
143
serde_json::to_value(&expected_after_first).unwrap()
144
serde_json::to_value(&expected_after_first).unwrap()
144
);
145
);
5 linesAll 18 lines5 lines
5 linesAll 18 lines5 lines
163
let expected_after_second: Vec<RolloutItem> = fork1_items[..cut_last_on_fork1].to_vec();
164
let expected_after_second: Vec<RolloutItem> = fork1_items[..cut_last_on_fork1].to_vec();
164
let fork2_items = read_items(&fork2_path);
165
let fork2_items = read_items(&fork2_path);
166
assert!(fork2_items.len() > expected_after_second.len());
165
pretty_assertions::assert_eq!(
167
pretty_assertions::assert_eq!(
166
serde_json::to_value(&fork2_items).unwrap(),
168
serde_json::to_value(&fork2_items[..expected_after_second.len()]).unwrap(),
167
serde_json::to_value(&expected_after_second).unwrap()
169
serde_json::to_value(&expected_after_second).unwrap()
168
);
170
);
169
}
171
}
prompt_caching.rscodex-rs/core/tests/suite
+85−82
Mark as viewed
5 linesAll 33 lines5 lines
5 linesAll 33 lines5 lines
34
fn default_env_context_str(cwd: &str, shell: &Shell) -> String {
34
fn default_env_context_str(cwd: &str, shell: &Shell) -> String {
35
let shell_name = shell.name();
35
let shell_name = shell.name();
36
format!(
36
format!(
37
r#"<environment_context>
37
r#"<environment_context>
38
<cwd>{cwd}</cwd>
38
<cwd>{cwd}</cwd>
39
<approval_policy>on-request</approval_policy>
40
<sandbox_mode>read-only</sandbox_mode>
41
<network_access>restricted</network_access>
42
<shell>{shell_name}</shell>
39
<shell>{shell_name}</shell>
43
</environment_context>"#
40
</environment_context>"#
44
)
41
)
45
}
42
}
5 linesAll 170 lines5 lines
5 linesAll 170 lines5 lines
216
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
213
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
217
async fn prefixes_context_and_instructions_once_and_consistently_across_requests()
214
async fn prefixes_context_and_instructions_once_and_consistently_across_requests()
218
-> anyhow::Result<()> {
215
-> anyhow::Result<()> {
5 linesAll 34 lines5 lines
5 linesAll 34 lines5 lines
253
let body1 = req1.single_request().body_json();
250
let body1 = req1.single_request().body_json();
254
let input1 = body1["input"].as_array().expect("input array");
251
let input1 = body1["input"].as_array().expect("input array");
255
assert_eq!(input1.len(), 3, "expected cached prefix + env + user msg");
252
assert_eq!(
253
input1.len(),
254
4,
InformationalR255-260
255
"expected permissions + cached prefix + env + user msg"
256
);
256
257
257
let ui_text = input1[0]["content"][0]["text"]
258
let ui_text = input1[1]["content"][0]["text"]
All 7 lines
All 7 lines
265
let shell = default_user_shell();
266
let shell = default_user_shell();
266
let cwd_str = config.cwd.to_string_lossy();
267
let cwd_str = config.cwd.to_string_lossy();
267
let expected_env_text = default_env_context_str(&cwd_str, &shell);
268
let expected_env_text = default_env_context_str(&cwd_str, &shell);
268
assert_eq!(
269
assert_eq!(
269
input1[1],
270
input1[2],
270
text_user_input(expected_env_text),
271
text_user_input(expected_env_text),
271
"expected environment context after UI message"
272
"expected environment context after UI message"
272
);
273
);
273
assert_eq!(input1[2], text_user_input("hello 1".to_string()));
274
assert_eq!(input1[3], text_user_input("hello 1".to_string()));
5 linesAll 10 lines5 lines
5 linesAll 10 lines5 lines
284
Ok(())
285
Ok(())
285
}
286
}
286
287
287
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
288
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
288
async fn overrides_turn_context_but_keeps_cached_prefix_and_key_constant() -> anyhow::Result<()> {
289
async fn overrides_turn_context_but_keeps_cached_prefix_and_key_constant() -> anyhow::Result<()> {
5 linesAll 25 lines5 lines
5 linesAll 25 lines5 lines
314
let writable = TempDir::new().unwrap();
315
let writable = TempDir::new().unwrap();
316
let new_policy = SandboxPolicy::WorkspaceWrite {
317
writable_roots: vec![writable.path().try_into().unwrap()],
318
network_access: true,
319
exclude_tmpdir_env_var: true,
320
exclude_slash_tmp: true,
321
};
315
codex
322
codex
316
.submit(Op::OverrideTurnContext {
323
.submit(Op::OverrideTurnContext {
317
cwd: None,
324
cwd: None,
318
approval_policy: Some(AskForApproval::Never),
325
approval_policy: Some(AskForApproval::Never),
319
sandbox_policy: Some(SandboxPolicy::WorkspaceWrite {
326
sandbox_policy: Some(new_policy.clone()),
320
writable_roots: vec![writable.path().try_into().unwrap()],
321
network_access: true,
322
exclude_tmpdir_env_var: true,
323
exclude_slash_tmp: true,
324
}),
325
model: Some("o3".to_string()),
327
model: Some("o3".to_string()),
326
effort: Some(Some(ReasoningEffort::High)),
328
effort: Some(Some(ReasoningEffort::High)),
327
summary: Some(ReasoningSummary::Detailed),
329
summary: Some(ReasoningSummary::Detailed),
328
})
330
})
329
.await?;
331
.await?;
5 linesAll 22 lines5 lines
5 linesAll 22 lines5 lines
352
let expected_user_message_2 = serde_json::json!({
354
let expected_user_message_2 = serde_json::json!({
353
"type": "message",
355
"type": "message",
354
"role": "user",
356
"role": "user",
355
"content": [ { "type": "input_text", "text": "hello 2" } ]
357
"content": [ { "type": "input_text", "text": "hello 2" } ]
356
});
358
});
357
// After overriding the turn context, the environment context should be emitted again
359
let expected_permissions_msg = body1["input"][0].clone();
358
// reflecting the new approval policy and sandbox settings. Omit cwd because it did
360
// After overriding the turn context, emit a new permissions message.
359
// not change.
361
let body1_input = body1["input"].as_array().expect("input array");
360
let shell = default_user_shell();
362
let expected_permissions_msg_2 = body2["input"][body1_input.len()].clone();
361
let expected_env_text_2 = format!(
363
assert_ne!(
362
r#"<environment_context>
364
expected_permissions_msg_2, expected_permissions_msg,
363
<approval_policy>never</approval_policy>
365
"expected updated permissions message after override"
364
<sandbox_mode>workspace-write</sandbox_mode>
365
<network_access>enabled</network_access>
366
<writable_roots>
367
<root>{}</root>
368
</writable_roots>
369
<shell>{}</shell>
370
</environment_context>"#,
371
writable.path().display(),
372
shell.name()
373
);
366
);
374
let expected_env_msg_2 = serde_json::json!({
367
let mut expected_body2 = body1["input"].as_array().expect("input array").to_vec();
375
"type": "message",
368
expected_body2.push(expected_permissions_msg_2);
376
"role": "user",
369
expected_body2.push(expected_user_message_2);
377
"content": [ { "type": "input_text", "text": expected_env_text_2 } ]
370
assert_eq!(body2["input"], serde_json::Value::Array(expected_body2));
378
});
379
let expected_body2 = serde_json::json!(
380
[
381 body1["input"].as_array().unwrap().as_slice(),
382 [expected_env_msg_2, expected_user_message_2].as_slice(),
383 ]
384
.concat()
385
);
386
assert_eq!(body2["input"], expected_body2);
387
371
388
Ok(())
372
Ok(())
389
}
373
}
390
374
391
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
375
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
392
async fn override_before_first_turn_emits_environment_context() -> anyhow::Result<()> {
376
async fn override_before_first_turn_emits_environment_context() -> anyhow::Result<()> {
5 linesAll 48 lines5 lines
5 linesAll 48 lines5 lines
441
assert!(
425
assert!(
442
env_texts
426
!env_texts.is_empty(),
443
.iter()
427
"expected environment context to be emitted: {env_texts:?}"
444
.any(|text| text.contains("<approval_policy>never</approval_policy>")),
445
"environment context should reflect overridden approval policy: {env_texts:?}"
446
);
428
);
447
429
448
let env_count = input
430
let env_count = input
449
.iter()
431
.iter()
450
.filter(|msg| {
432
.filter(|msg| {
5 linesAll 12 lines5 lines
5 linesAll 12 lines5 lines
463
})
445
})
464
.count();
446
.count();
465
assert_eq!(
447
assert!(
466
env_count, 2,
448
env_count >= 1,
467
"environment context should appear exactly twice, found {env_count}"
449
"environment context should appear at least once, found {env_count}"
450
);
451
452
let permissions_texts: Vec<&str> = input
453
.iter()
454
.filter_map(|msg| {
455
let role = msg["role"].as_str()?;
456
if role != "developer" {
457
return None;
458
}
459
msg["content"]
460
.as_array()
461
.and_then(|content| content.first())
462
.and_then(|item| item["text"].as_str())
463
})
464
.collect();
465
assert!(
466
permissions_texts
467
.iter()
468
.any(|text| text.contains("`approval_policy` is `never`")),
469
"permissions message should reflect overridden approval policy: {permissions_texts:?}"
468
);
470
);
5 linesAll 16 lines5 lines
5 linesAll 16 lines5 lines
485
}
487
}
486
488
487
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
489
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
488
async fn per_turn_overrides_keep_cached_prefix_and_key_constant() -> anyhow::Result<()> {
490
async fn per_turn_overrides_keep_cached_prefix_and_key_constant() -> anyhow::Result<()> {
5 linesAll 27 lines5 lines
5 linesAll 27 lines5 lines
516
let writable = TempDir::new().unwrap();
518
let writable = TempDir::new().unwrap();
519
let new_policy = SandboxPolicy::WorkspaceWrite {
520
writable_roots: vec![AbsolutePathBuf::try_from(writable.path()).unwrap()],
521
network_access: true,
522
exclude_tmpdir_env_var: true,
523
exclude_slash_tmp: true,
524
};
517
codex
525
codex
518
.submit(Op::UserTurn {
526
.submit(Op::UserTurn {
All 5 lines
All 5 lines
524
sandbox_policy: SandboxPolicy::WorkspaceWrite {
532
sandbox_policy: new_policy.clone(),
525
writable_roots: vec![AbsolutePathBuf::try_from(writable.path()).unwrap()],
526
network_access: true,
527
exclude_tmpdir_env_var: true,
528
exclude_slash_tmp: true,
529
},
All 4 lines
All 4 lines
534
})
537
})
535
.await?;
538
.await?;
5 linesAll 20 lines5 lines
5 linesAll 20 lines5 lines
556
let expected_env_text_2 = format!(
559
let expected_env_text_2 = format!(
557
r#"<environment_context>
560
r#"<environment_context>
558
<cwd>{}</cwd>
561
<cwd>{}</cwd>
559
<approval_policy>never</approval_policy>
560
<sandbox_mode>workspace-write</sandbox_mode>
561
<network_access>enabled</network_access>
562
<writable_roots>
563
<root>{}</root>
564
</writable_roots>
565
<shell>{}</shell>
562
<shell>{}</shell>
566
</environment_context>"#,
563
</environment_context>"#,
567
new_cwd.path().display(),
564
new_cwd.path().display(),
568
writable.path().display(),
565
shell.name()
569
shell.name(),
570
);
566
);
All 5 lines
All 5 lines
576
let expected_body2 = serde_json::json!(
572
let expected_permissions_msg = body1["input"][0].clone();
577
[
573 let body1_input = body1["input"].as_array().expect("input array");
578 body1["input"].as_array().unwrap().as_slice(),
574 let expected_permissions_msg_2 = body2["input"][body1_input.len() + 1].clone();
579 [expected_env_msg_2, expected_user_message_2].as_slice(),
575 assert_ne!(
580 ]
576
expected_permissions_msg_2, expected_permissions_msg,
581
.concat()
577
"expected updated permissions message after per-turn override"
582
);
578
);
583
assert_eq!(body2["input"], expected_body2);
579
let mut expected_body2 = body1_input.to_vec();
580
expected_body2.push(expected_env_msg_2);
581
expected_body2.push(expected_permissions_msg_2);
582
expected_body2.push(expected_user_message_2);
583
assert_eq!(body2["input"], serde_json::Value::Array(expected_body2));
584
584
585
Ok(())
585
Ok(())
586
}
586
}
587
587
588
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
588
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
589
async fn send_user_turn_with_no_changes_does_not_send_environment_context() -> anyhow::Result<()> {
589
async fn send_user_turn_with_no_changes_does_not_send_environment_context() -> anyhow::Result<()> {
5 linesAll 61 lines5 lines
5 linesAll 61 lines5 lines
651
let expected_ui_msg = body1["input"][0].clone();
651
let expected_permissions_msg = body1["input"][0].clone();
652
let expected_ui_msg = body1["input"][1].clone();
All 7 lines
All 7 lines
659
let expected_input_1 = serde_json::Value::Array(vec![
660 let expected_input_1 = serde_json::Value::Array(vec![
661 expected_permissions_msg.clone(),
660 expected_ui_msg.clone(),
662 expected_ui_msg.clone(),
661 expected_env_msg_1.clone(),
663 expected_env_msg_1.clone(),
662 expected_user_message_1.clone(),
664 expected_user_message_1.clone(),
663 ]);
665 ]);
664
assert_eq!(body1["input"], expected_input_1);
666
assert_eq!(body1["input"], expected_input_1);
665
667
666
let expected_user_message_2 = text_user_input("hello 2".to_string());
668
let expected_user_message_2 = text_user_input("hello 2".to_string());
667
let expected_input_2 = serde_json::Value::Array(vec![
669 let expected_input_2 = serde_json::Value::Array(vec![
670 expected_permissions_msg,
668 expected_ui_msg,
671 expected_ui_msg,
669 expected_env_msg_1,
672 expected_env_msg_1,
670 expected_user_message_1,
673 expected_user_message_1,
671 expected_user_message_2,
674 expected_user_message_2,
672 ]);
675 ]);
673
assert_eq!(body2["input"], expected_input_2);
676
assert_eq!(body2["input"], expected_input_2);
674
677
675
Ok(())
678
Ok(())
676
}
679
}
677
680
678
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
681
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
679
async fn send_user_turn_with_changes_sends_environment_context() -> anyhow::Result<()> {
682
async fn send_user_turn_with_changes_sends_environment_context() -> anyhow::Result<()> {
5 linesAll 61 lines5 lines
5 linesAll 61 lines5 lines
741
let expected_ui_msg = body1["input"][0].clone();
744
let expected_permissions_msg = body1["input"][0].clone();
745
let expected_ui_msg = body1["input"][1].clone();
All 5 lines
All 5 lines
747
let expected_input_1 = serde_json::Value::Array(vec![
751 let expected_input_1 = serde_json::Value::Array(vec![
752 expected_permissions_msg.clone(),
748 expected_ui_msg.clone(),
753 expected_ui_msg.clone(),
749 expected_env_msg_1.clone(),
754 expected_env_msg_1.clone(),
750 expected_user_message_1.clone(),
755 expected_user_message_1.clone(),
751 ]);
756 ]);
752
assert_eq!(body1["input"], expected_input_1);
757
assert_eq!(body1["input"], expected_input_1);
753
758
754
let shell_name = shell.name();
759
let body1_input = body1["input"].as_array().expect("input array");
755
let expected_env_msg_2 = text_user_input(format!(
760
let expected_permissions_msg_2 = body2["input"][body1_input.len()].clone();
756
r#"<environment_context>
761
assert_ne!(
757
<approval_policy>never</approval_policy>
762
expected_permissions_msg_2, expected_permissions_msg,
758
<sandbox_mode>danger-full-access</sandbox_mode>
763
"expected updated permissions message after policy change"
759
<network_access>enabled</network_access>
764
);
760
<shell>{shell_name}</shell>
761
</environment_context>"#
762
));
763
let expected_user_message_2 = text_user_input("hello 2".to_string());
765
let expected_user_message_2 = text_user_input("hello 2".to_string());
764
let expected_input_2 = serde_json::Value::Array(vec![
766 let expected_input_2 = serde_json::Value::Array(vec![
767 expected_permissions_msg,
765 expected_ui_msg,
768 expected_ui_msg,
766 expected_env_msg_1,
769 expected_env_msg_1,
767 expected_user_message_1,
770 expected_user_message_1,
768 expected_env_msg_2,
771 expected_permissions_msg_2,
769 expected_user_message_2,
772 expected_user_message_2,
770 ]);
773 ]);
771
assert_eq!(body2["input"], expected_input_2);
774
assert_eq!(body2["input"], expected_input_2);
772
775
773
Ok(())
776
Ok(())
774
}
777
}
codex_tool.rscodex-rs/mcp-server/tests/suite
+17−14
Mark as viewed
5 linesAll 332 lines5 lines
5 linesAll 332 lines5 lines
333
async fn codex_tool_passes_base_instructions() -> anyhow::Result<()> {
333
async fn codex_tool_passes_base_instructions() -> anyhow::Result<()> {
5 linesAll 45 lines5 lines
5 linesAll 45 lines5 lines
379
let requests = server.received_requests().await.unwrap();
379
let requests = server.received_requests().await.unwrap();
380
let request = requests[0].body_json::<serde_json::Value>()?;
380
let request = requests[0].body_json::<serde_json::Value>()?;
381
let instructions = request["messages"][0]["content"].as_str().unwrap();
381
let instructions = request["messages"][0]["content"].as_str().unwrap();
382
assert!(instructions.starts_with("You are a helpful assistant."));
382
assert!(instructions.starts_with("You are a helpful assistant."));
383
383
384
let developer_msg = request["messages"]
384
let developer_messages: Vec<&serde_json::Value> = request["messages"]
385
.as_array()
385
.as_array()
386
.and_then(|messages| {
386
.unwrap()
387
messages
387
.iter()
388
.iter()
388
.filter(|msg| msg.get("role").and_then(|role| role.as_str()) == Some("developer"))
389
.find(|msg| msg.get("role").and_then(|role| role.as_str()) == Some("developer"))
389
.collect();
390
})
390
let developer_contents: Vec<&str> = developer_messages
391
.unwrap();
391
.iter()
392
let developer_content = developer_msg
392
.filter_map(|msg| msg.get("content").and_then(|value| value.as_str()))
393
.get("content")
393
.collect();
394
.and_then(|value| value.as_str())
394
assert!(
395
.unwrap();
395
developer_contents
396
.iter()
397
.any(|content| content.contains("`sandbox_mode`")),
398
"expected permissions developer message, got {developer_contents:?}"
399
);
396
assert!(
400
assert!(
397
!developer_content.contains('<'),
401
developer_contents.contains(&"Foreshadow upcoming tool calls."),
398
"expected developer instructions without XML tags, got `{developer_content}`"
402
"expected developer instructions in developer messages, got {developer_contents:?}"
399
);
403
);
400
assert_eq!(developer_content, "Foreshadow upcoming tool calls.");
401
404
402
Ok(())
405
Ok(())
403
}
406
}
5 linesAll 87 lines5 lines
5 linesAll 87 lines5 lines
End of changesChat about this PR
InfoChat
1 Potential bug
Permissions message not updated when cwd changes in WorkspaceWrite mode
Bugcodex.rs:1017
5 Flags
Resume/fork intentionally duplicates initial context for policy synchronization
codex.rs:856-859
EnvironmentContext simplified - sandbox info moved to permissions message
environment_context.rs:12-16
Prompt templates use placeholder interpolation for network_access
models.rs:272-281
Test expectations updated for new message ordering
prompt_caching.rs:255-260
Static prompts removed from markdown - no longer bundled in instructions
prompt.md
ChecksPassed33/33
All checks passed
Reviewers3
DYdylan-hurd-oai
CHchatgpt-codex-connector
PApakrym-oai
Assignees
No assignees
Labels
No labels assigned
Introducing Devin Review
Devin Review is an intelligent PR review tool that helps you better understand your team's code before you ship it.
1
Intelligently organizes your diff
Groups changes into logical sections with clear explanations
→
2
Detects moved or copied code
Keeps the diff clean by highlighting copy-pasted or moved code
3
Analyzes your PR
Highlights potential bugs and flags areas that deserve extra careful review
Ask questions or request edits in the chat.
Ask Devin anything about this PR (Ctrl+I)
Authorization Response