Release v1.48.0 - #5079
Release v1.48.0#5079
Conversation
Co-authored-by: openhands <openhands@all-hands.dev>
|
Hi! I started running the integration tests on your PR. You will receive a comment with the results shortly. |
|
Hi! I started running the behavior tests on your PR. You will receive a comment with the results shortly. |
|
👋 This PR needs a couple of things fixed before OpenHands can review it:
Push an update once this is addressed and this check re-runs automatically. This is an automated check - no AI was used to generate this comment. |
REST API breakage checks (OpenAPI) — ✅ PASSEDResult: ✅ PASSED |
🔒 Release Security Scan🔒 Approval drift (time-of-check vs time-of-use)❌ 8 finding(s) Baseline:
Audited 49 PR(s): 41 clean, 8 flagged, 0 un-auditable. 📦 Supply-chain dependency diff✅ no findings Baseline: OSV: no known vulns across 20 new/bumped dep(s). Added dependencies
Bumped dependencies
Internal `openhands-*` bumps
Deterministic scanners: approval-drift + supply-chain dependency diff. Read-only; no PR code executed. |
🧪 Integration Tests ResultsOverall Success Rate: 77.3% 📁 Detailed Logs & ArtifactsClick the links below to access detailed agent/LLM logs showing the complete reasoning process for each model. On the GitHub Actions page, scroll down to the 'Artifacts' section to download the logs.
📊 Summary
📋 Detailed Resultslitellm_proxy_minimax_MiniMax_M2.7
Skipped Tests:
litellm_proxy_gemini_3.1_pro_preview
litellm_proxy_deepseek_deepseek_v4_flash
litellm_proxy_openai_gpt_5.5
Failed Tests:
litellm_proxy_anthropic_claude_sonnet_4_6
Failed Tests:
|
Coverage Report •
|
||||||||||||||||||||||||||||||
🔄 Running Examples with
|
| Example | Status | Duration | Cost |
|---|---|---|---|
| 01_standalone_sdk/02_custom_tools.py | ✅ PASS | 26.3s | $0.03 |
| 01_standalone_sdk/03_activate_skill.py | ✅ PASS | 23.6s | $0.03 |
| 01_standalone_sdk/05_use_llm_registry.py | ✅ PASS | 12.5s | $0.01 |
| 01_standalone_sdk/07_mcp_integration.py | ✅ PASS | 42.0s | $0.03 |
| 01_standalone_sdk/09_pause_example.py | ✅ PASS | 14.9s | $0.01 |
| 01_standalone_sdk/10_persistence.py | ✅ PASS | 23.9s | $0.02 |
| 01_standalone_sdk/11_async.py | ✅ PASS | 38.7s | $0.04 |
| 01_standalone_sdk/12_custom_secrets.py | ✅ PASS | 14.4s | $0.01 |
| 01_standalone_sdk/13_get_llm_metrics.py | ✅ PASS | 35.0s | $0.03 |
| 01_standalone_sdk/14_context_condenser.py | ✅ PASS | 3m 37s | $0.23 |
| 01_standalone_sdk/17_image_input.py | ✅ PASS | 28.3s | $0.02 |
| 01_standalone_sdk/18_send_message_while_processing.py | ✅ PASS | 28.5s | $0.02 |
| 01_standalone_sdk/19_llm_routing.py | ✅ PASS | 18.3s | $0.02 |
| 01_standalone_sdk/20_stuck_detector.py | ✅ PASS | 22.5s | $0.03 |
| 01_standalone_sdk/21_generate_extraneous_conversation_costs.py | ✅ PASS | 12.1s | $0.00 |
| 01_standalone_sdk/22_anthropic_thinking.py | ✅ PASS | 15.6s | $0.01 |
| 01_standalone_sdk/23_responses_reasoning.py | ❌ FAIL Exit code 1 |
7.1s | -- |
| 01_standalone_sdk/24_planning_agent_workflow.py | ✅ PASS | 4m 36s | $0.25 |
| 01_standalone_sdk/25_agent_delegation.py | ✅ PASS | 1m 6s | $0.06 |
| 01_standalone_sdk/26_custom_visualizer.py | ✅ PASS | 28.6s | $0.03 |
| 01_standalone_sdk/28_ask_agent_example.py | ❌ FAIL Exit code 1 |
9.1s | -- |
| 01_standalone_sdk/29_llm_streaming.py | ✅ PASS | 33.8s | $0.02 |
| 01_standalone_sdk/30_tom_agent.py | ✅ PASS | 11.3s | $0.01 |
| 01_standalone_sdk/31_iterative_refinement.py | ✅ PASS | 4m 40s | $0.25 |
| 01_standalone_sdk/32_configurable_security_policy.py | ✅ PASS | 21.0s | $0.02 |
| 01_standalone_sdk/33_hooks/main.py | ✅ PASS | 37.2s | $0.04 |
| 01_standalone_sdk/34_critic_example.py | ✅ PASS | 2m 35s | $0.03 |
| 01_standalone_sdk/36_event_json_to_openai_messages.py | ✅ PASS | 10.6s | $0.00 |
| 01_standalone_sdk/37_llm_profile_store/main.py | ✅ PASS | 6.3s | $0.00 |
| 01_standalone_sdk/38_browser_session_recording.py | ✅ PASS | 41.0s | $0.03 |
| 01_standalone_sdk/39_llm_fallback.py | ✅ PASS | 12.9s | $0.01 |
| 01_standalone_sdk/40_acp_agent_example.py | ✅ PASS | 47.4s | $0.33 |
| 01_standalone_sdk/41_task_tool_set.py | ✅ PASS | 27.4s | $0.03 |
| 01_standalone_sdk/42_file_based_subagents.py | ✅ PASS | 53.5s | $0.05 |
| 01_standalone_sdk/44_model_switching_in_convo.py | ✅ PASS | 10.7s | $0.01 |
| 01_standalone_sdk/45_parallel_tool_execution.py | ✅ PASS | 3m 45s | $0.46 |
| 01_standalone_sdk/46_agent_settings.py | ✅ PASS | 13.7s | $0.01 |
| 01_standalone_sdk/47_defense_in_depth_security.py | ✅ PASS | 4.3s | $0.00 |
| 01_standalone_sdk/48_conversation_fork.py | ✅ PASS | 28.0s | $0.01 |
| 01_standalone_sdk/49_switch_llm_tool.py | ✅ PASS | 10.7s | $0.04 |
| 01_standalone_sdk/50_async_cancellation.py | ✅ PASS | 14.6s | $0.01 |
| 01_standalone_sdk/51_agent_hooks/main.py | ✅ PASS | 28.3s | $0.04 |
| 01_standalone_sdk/52_dynamic_workflow.py | ✅ PASS | 2m 36s | $0.09 |
| 01_standalone_sdk/53_client_defined_tools.py | ✅ PASS | 16.4s | $0.01 |
| 01_standalone_sdk/54_goal_completion_loop.py | ✅ PASS | 40.5s | $0.03 |
| 01_standalone_sdk/55_persistent_memory.py | ✅ PASS | 18.9s | $0.02 |
| 01_standalone_sdk/56_structured_output.py | ✅ PASS | 28.7s | $0.04 |
| 01_standalone_sdk/57_prompt_hooks/main.py | ✅ PASS | 13.3s | $0.00 |
| 01_standalone_sdk/58_ask_oracle_tool/main.py | ✅ PASS | 15.6s | $0.01 |
| 02_remote_agent_server/01_convo_with_local_agent_server.py | ✅ PASS | 40.4s | $0.02 |
| 02_remote_agent_server/02_convo_with_docker_sandboxed_server.py | ✅ PASS | 1m 34s | $0.04 |
| 02_remote_agent_server/03_browser_use_with_docker_sandboxed_server.py | ✅ PASS | 1m 39s | $0.08 |
| 02_remote_agent_server/04_convo_with_api_sandboxed_server.py | ✅ PASS | 1m 49s | $0.05 |
| 02_remote_agent_server/06_custom_tool/main.py | ✅ PASS | 5m 10s | $0.03 |
| 02_remote_agent_server/07_convo_with_cloud_workspace.py | ✅ PASS | 1m 26s | $0.04 |
| 02_remote_agent_server/08_convo_with_apptainer_sandboxed_server.py | ✅ PASS | 3m 58s | $0.02 |
| 02_remote_agent_server/09_acp_agent_with_remote_runtime.py | ✅ PASS | 1m 8s | $0.39 |
| 02_remote_agent_server/10_cloud_workspace_share_credentials.py | ✅ PASS | 1m 18s | $0.00 |
| 02_remote_agent_server/11_conversation_fork.py | ✅ PASS | 49.8s | $0.00 |
| 02_remote_agent_server/12_settings_and_secrets_api.py | ✅ PASS | 2m 20s | $0.01 |
| 02_remote_agent_server/13_workspace_get_llm.py | ✅ PASS | 30.0s | $0.01 |
| 02_remote_agent_server/14_client_defined_tools.py | ✅ PASS | 32.8s | $0.02 |
| 02_remote_agent_server/15_openai_compatible_gateway.py | ✅ PASS | 24.7s | $0.01 |
| 02_remote_agent_server/16_deferred_init.py | ✅ PASS | 15.3s | $0.01 |
| 02_remote_agent_server/17_convo_with_agent_sandbox_server.py | ❌ FAIL Exit code 1 |
6.3s | -- |
| 04_llm_specific_tools/01_gpt5_apply_patch_preset.py | ❌ FAIL Exit code 1 |
13.4s | -- |
| 04_llm_specific_tools/02_gemini_file_tools.py | ✅ PASS | 43.1s | $0.10 |
| 05_skills_and_plugins/01_loading_agentskills/main.py | ✅ PASS | 18.0s | $0.02 |
| 05_skills_and_plugins/02_loading_plugins/main.py | ✅ PASS | 23.2s | $0.03 |
| 05_skills_and_plugins/04_mixed_marketplace_skills/main.py | ✅ PASS | 6.7s | $0.00 |
❌ Some tests failed
Total: 70 | Passed: 66 | Failed: 4 | Total Cost: $3.36
Failed examples:
- examples/01_standalone_sdk/23_responses_reasoning.py: Exit code 1
- examples/01_standalone_sdk/28_ask_agent_example.py: Exit code 1
- examples/02_remote_agent_server/17_convo_with_agent_sandbox_server.py: Exit code 1
- examples/04_llm_specific_tools/01_gpt5_apply_patch_preset.py: Exit code 1
🧪 Integration Tests ResultsOverall Success Rate: 80.0% 📁 Detailed Logs & ArtifactsClick the links below to access detailed agent/LLM logs showing the complete reasoning process for each model. On the GitHub Actions page, scroll down to the 'Artifacts' section to download the logs.
📊 Summary
📋 Detailed Resultslitellm_proxy_minimax_MiniMax_M2.7
litellm_proxy_gemini_3.1_pro_preview
litellm_proxy_deepseek_deepseek_v4_flash
litellm_proxy_openai_gpt_5.5
Failed Tests:
litellm_proxy_anthropic_claude_sonnet_4_6
|
|
I went through the security scan and I didn’t spot messed up stuff / supply chain / weird deps or something. I didn’t use my agent, out of principle, but I think maybe next time I can do this scan with a SOTA LLM like Fable 5.1. For better or worse, Anthropic seems to have fairly tough classifiers on it… and human attention, even human skimming, will not scale with diffs as large as these. |
🔄 Running Examples with
|
| Example | Status | Duration | Cost |
|---|---|---|---|
| 01_standalone_sdk/02_custom_tools.py | ✅ PASS | 22.3s | $0.03 |
| 01_standalone_sdk/03_activate_skill.py | ✅ PASS | 23.4s | $0.03 |
| 01_standalone_sdk/05_use_llm_registry.py | ✅ PASS | 10.4s | $0.01 |
| 01_standalone_sdk/07_mcp_integration.py | ✅ PASS | 34.2s | $0.03 |
| 01_standalone_sdk/09_pause_example.py | ✅ PASS | 11.2s | $0.01 |
| 01_standalone_sdk/10_persistence.py | ✅ PASS | 25.6s | $0.02 |
| 01_standalone_sdk/11_async.py | ✅ PASS | 33.4s | $0.04 |
| 01_standalone_sdk/12_custom_secrets.py | ✅ PASS | 13.7s | $0.01 |
| 01_standalone_sdk/13_get_llm_metrics.py | ✅ PASS | 31.1s | $0.04 |
| 01_standalone_sdk/14_context_condenser.py | ✅ PASS | 2m 43s | $0.16 |
| 01_standalone_sdk/17_image_input.py | ✅ PASS | 22.1s | $0.02 |
| 01_standalone_sdk/18_send_message_while_processing.py | ✅ PASS | 24.7s | $0.02 |
| 01_standalone_sdk/19_llm_routing.py | ✅ PASS | 18.7s | $0.02 |
| 01_standalone_sdk/20_stuck_detector.py | ✅ PASS | 16.1s | $0.02 |
| 01_standalone_sdk/21_generate_extraneous_conversation_costs.py | ✅ PASS | 11.1s | $0.00 |
| 01_standalone_sdk/22_anthropic_thinking.py | ✅ PASS | 16.9s | $0.01 |
| 01_standalone_sdk/23_responses_reasoning.py | ❌ FAIL Exit code 1 |
47.0s | -- |
| 01_standalone_sdk/24_planning_agent_workflow.py | ✅ PASS | 4m 20s | $0.31 |
| 01_standalone_sdk/25_agent_delegation.py | ✅ PASS | 56.1s | $0.05 |
| 01_standalone_sdk/26_custom_visualizer.py | ✅ PASS | 20.5s | $0.02 |
| 01_standalone_sdk/28_ask_agent_example.py | ✅ PASS | 32.0s | $0.05 |
| 01_standalone_sdk/29_llm_streaming.py | ✅ PASS | 33.2s | $0.02 |
| 01_standalone_sdk/30_tom_agent.py | ✅ PASS | 14.9s | $0.02 |
| 01_standalone_sdk/31_iterative_refinement.py | ✅ PASS | 4m 38s | $0.26 |
| 01_standalone_sdk/32_configurable_security_policy.py | ✅ PASS | 17.0s | $0.02 |
| 01_standalone_sdk/33_hooks/main.py | ✅ PASS | 33.8s | $0.04 |
| 01_standalone_sdk/34_critic_example.py | ✅ PASS | 1m 57s | $0.02 |
| 01_standalone_sdk/36_event_json_to_openai_messages.py | ✅ PASS | 11.9s | $0.00 |
| 01_standalone_sdk/37_llm_profile_store/main.py | ✅ PASS | 9.8s | $0.00 |
| 01_standalone_sdk/38_browser_session_recording.py | ✅ PASS | 43.4s | $0.04 |
| 01_standalone_sdk/39_llm_fallback.py | ✅ PASS | 25.9s | $0.01 |
| 01_standalone_sdk/40_acp_agent_example.py | ✅ PASS | 51.0s | $0.33 |
| 01_standalone_sdk/41_task_tool_set.py | ✅ PASS | 24.2s | $0.03 |
| 01_standalone_sdk/42_file_based_subagents.py | ✅ PASS | 50.9s | $0.05 |
| 01_standalone_sdk/44_model_switching_in_convo.py | ✅ PASS | 10.0s | $0.01 |
| 01_standalone_sdk/45_parallel_tool_execution.py | ✅ PASS | 5m 14s | $0.54 |
| 01_standalone_sdk/46_agent_settings.py | ✅ PASS | 13.4s | $0.01 |
| 01_standalone_sdk/47_defense_in_depth_security.py | ✅ PASS | 4.6s | $0.00 |
| 01_standalone_sdk/48_conversation_fork.py | ✅ PASS | 22.8s | $0.01 |
| 01_standalone_sdk/49_switch_llm_tool.py | ❌ FAIL Exit code 1 |
3.9s | -- |
| 01_standalone_sdk/50_async_cancellation.py | ✅ PASS | 13.9s | $0.00 |
| 01_standalone_sdk/51_agent_hooks/main.py | ✅ PASS | 53.2s | $0.06 |
| 01_standalone_sdk/52_dynamic_workflow.py | ✅ PASS | 4m 26s | $0.16 |
| 01_standalone_sdk/53_client_defined_tools.py | ✅ PASS | 14.8s | $0.01 |
| 01_standalone_sdk/54_goal_completion_loop.py | ✅ PASS | 34.5s | $0.03 |
| 01_standalone_sdk/55_persistent_memory.py | ✅ PASS | 17.0s | $0.02 |
| 01_standalone_sdk/56_structured_output.py | ✅ PASS | 42.3s | $0.04 |
| 01_standalone_sdk/57_prompt_hooks/main.py | ✅ PASS | 20.5s | $0.00 |
| 01_standalone_sdk/58_ask_oracle_tool/main.py | ✅ PASS | 18.6s | $0.01 |
| 02_remote_agent_server/01_convo_with_local_agent_server.py | ✅ PASS | 39.5s | $0.02 |
| 02_remote_agent_server/02_convo_with_docker_sandboxed_server.py | ✅ PASS | 1m 34s | $0.04 |
| 02_remote_agent_server/03_browser_use_with_docker_sandboxed_server.py | ✅ PASS | 1m 50s | $0.13 |
| 02_remote_agent_server/04_convo_with_api_sandboxed_server.py | ✅ PASS | 1m 46s | $0.03 |
| 02_remote_agent_server/06_custom_tool/main.py | ✅ PASS | 5m 19s | $0.04 |
| 02_remote_agent_server/07_convo_with_cloud_workspace.py | ✅ PASS | 1m 9s | $0.03 |
| 02_remote_agent_server/08_convo_with_apptainer_sandboxed_server.py | ✅ PASS | 3m 20s | $0.02 |
| 02_remote_agent_server/09_acp_agent_with_remote_runtime.py | ✅ PASS | 1m 43s | $0.45 |
| 02_remote_agent_server/10_cloud_workspace_share_credentials.py | ✅ PASS | 1m 6s | $0.00 |
| 02_remote_agent_server/11_conversation_fork.py | ✅ PASS | 41.1s | $0.00 |
| 02_remote_agent_server/12_settings_and_secrets_api.py | ✅ PASS | 2m 11s | $0.01 |
| 02_remote_agent_server/13_workspace_get_llm.py | ✅ PASS | 31.8s | $0.01 |
| 02_remote_agent_server/14_client_defined_tools.py | ✅ PASS | 27.7s | $0.02 |
| 02_remote_agent_server/15_openai_compatible_gateway.py | ✅ PASS | 22.4s | $0.01 |
| 02_remote_agent_server/16_deferred_init.py | ✅ PASS | 22.0s | $0.01 |
| 04_llm_specific_tools/01_gpt5_apply_patch_preset.py | ❌ FAIL Exit code 1 |
8.7s | -- |
| 04_llm_specific_tools/02_gemini_file_tools.py | ✅ PASS | 36.4s | $0.08 |
| 05_skills_and_plugins/01_loading_agentskills/main.py | ✅ PASS | 16.5s | $0.02 |
| 05_skills_and_plugins/02_loading_plugins/main.py | ✅ PASS | 21.7s | $0.03 |
| 05_skills_and_plugins/04_mixed_marketplace_skills/main.py | ✅ PASS | 3.8s | $0.00 |
❌ Some tests failed
Total: 69 | Passed: 66 | Failed: 3 | Total Cost: $3.56
Failed examples:
- examples/01_standalone_sdk/23_responses_reasoning.py: Exit code 1
- examples/01_standalone_sdk/49_switch_llm_tool.py: Exit code 1
- examples/04_llm_specific_tools/01_gpt5_apply_patch_preset.py: Exit code 1
Release v1.48.0
This PR prepares the release for version 1.48.0.
Started by: @VascoSch92
Release Checklist
integration-test)behavior-test)test-examples)security-scan)release-note-requiredPRs are accurately called out in the final release notesWhat happens on merge
When this PR is merged, the
create-release.ymlworkflow will automatically:v1.48.0and auto-generated notes, plus an explicit preamble for mergedrelease-note-requiredPRsv1.48.0version-bump-prs.ymlafter successful PyPI publication🐳 Agent Server images for this PR — GHCR package, pull/run commands, and all pushed tags (click to expand)
• GHCR package: https://github.com/OpenHands/agent-sdk/pkgs/container/agent-server
Variants & Base Images
eclipse-temurin:17-jdkpython-node-runtimepython-node-runtimegolang:1.21-bookwormPull (multi-arch manifest)
# Each variant is a multi-arch manifest supporting both amd64 and arm64 docker pull ghcr.io/openhands/agent-server:7d44742-pythonRun
All tags pushed for this build
About Multi-Architecture Support
7d44742-python) is a multi-arch manifest supporting both amd64 and arm647d44742-python-amd64) are also available if needed