Skip to content

feat: Endpoint Persistence using Network Volume (phase 1) - #25

Merged
deanq merged 90 commits into
mainfrom
deanq/ae-1092-tetra-volume-warm-cache
Oct 6, 2025
Merged

feat: Endpoint Persistence using Network Volume (phase 1)#25
deanq merged 90 commits into
mainfrom
deanq/ae-1092-tetra-volume-warm-cache

Conversation

@deanq

@deanq deanq commented Aug 22, 2025

Copy link
Copy Markdown
Collaborator

Phase 1: Remove workspace runtimes from network volume. Simplify any download acceleration that is affected by this refactor. This should speed up any of our examples that have network volumes attached.

Phase 2: Implement the volume cache replicator (#31)

deanq added 30 commits August 15, 2025 17:05
Add core download acceleration modules with aria2c integration:
- download_accelerator.py: Main acceleration classes with multi-connection downloads
- huggingface_accelerator.py: Specialized HF model acceleration
- constants.py: Download acceleration configuration constants
- __init__.py: Package structure for src module
Enhanced dependency installation with intelligent acceleration:
- Auto-detects large packages for acceleration (torch, transformers, etc.)
- Integrates with remote executor for acceleration control
- Maintains backward compatibility with existing workflows
- Provides graceful fallback when aria2c unavailable
Enhanced workspace manager with HuggingFace model pre-caching:
- Pre-cache specified HF models before function execution
- Integrates with volume-aware caching system
- Optimizes cold start times for ML workloads
Comprehensive test suite for download acceleration:
- Integration tests for aria2 detection and fallback behavior
- HF model acceleration testing with authentication
- Volume-aware acceleration scenarios
- Error handling and performance validation
- Update test files moved to src/ directory
- Enhanced test coverage for acceleration features
- Updated dependencies and documentation
- Submodule updates for tetra-rp
- Added nala accelerated installation for large system packages
- Enhanced DependencyInstaller with automatic nala fallback to apt-get
- Updated Docker images to include nala package manager
- Added comprehensive system package acceleration tests
- Improved acceleration logging with system package status
Simplify dependency installation by removing aria2c acceleration for Python packages.
UV's built-in parallel downloading and caching is superior and eliminates the need
for additional complexity.

Changes:
- Remove LARGE_PACKAGE_PATTERNS from constants.py
- Simplify DependencyInstaller.install_dependencies() to single parameter
- Remove Python package acceleration logic and related methods
- Update RemoteExecutor to use simplified API
- Update tests to match new simplified interface

System package acceleration (nala) and HuggingFace model acceleration remain intact
as they provide meaningful performance benefits over standard tools.

Core functionality verified:
- All handler tests pass (8/8)
- All unit tests pass (98/98)
- Code quality checks pass (format, lint, typecheck)
Add conditional acceleration logic - passes accelerate_downloads to installers, HF model caching only when accelerated + models specified
…bled

Implement _install_with_pip() method and route between UV (accelerated) vs pip (standard) based on accelerate_downloads parameter
Add HfXetDownloader for subsequent downloads, implement smart strategy: hf_xet for cached files → hf_transfer for fresh downloads → fallback
Add tests for both acceleration enabled/disabled scenarios, verify UV vs pip routing, update existing test assertions
Update test expectations to handle accelerate_downloads parameter in integration scenarios
Update build files and dependency locks to support new acceleration functionality
Always use UV for Python package installation regardless of acceleration setting.
The _install_with_pip method has been removed as UV provides more reliable
virtual environment handling and package management.

- Remove _install_with_pip() method (70 lines)
- Simplify install_dependencies() to always use UV
- Maintain differential installation when acceleration is enabled
Update dependency installer tests to reflect the removal of pip support:
- Fix test_install_dependencies_with_acceleration_disabled to expect UV
- Rename test_install_dependencies_pip_failure to test_install_dependencies_uv_failure
- Update assertions to check for "uv pip" commands
- Update test descriptions and expected error messages

All tests now correctly validate UV-only package installation behavior.
Rename test_pip_no_acceleration.json to test_uv_no_acceleration.json
and update content to reflect UV-only package installation:
- Update function name from test_pip_installation_without_acceleration
  to test_uv_installation_without_acceleration
- Update success message to reference UV instead of pip
- Maintain same test logic for package import validation

This test validates that packages installed with accelerate_downloads=False
are properly available using UV package manager.
Add parallel installation of dependencies when acceleration is enabled:
- Add async wrappers for dependency and model download methods
- Implement _install_dependencies_parallel() using asyncio.gather()
- Add _install_dependencies_sequential() for non-accelerated path
- Add _process_parallel_results() for error handling
- Route between parallel/sequential execution based on accelerate_downloads flag

When accelerate_downloads=True, system packages, Python packages, and HF model
downloads execute concurrently for improved performance.
Add accelerate_model_download_async() method to WorkspaceManager to support
parallel execution of model downloads when acceleration is enabled.

This async wrapper allows HF model downloads to run concurrently with
dependency installations for improved performance.
Update test mocks and expectations for parallel execution implementation:
- Fix AsyncMock setup for async dependency installation methods
- Update test_dependency_management.py for async method calls
- Update test_download_acceleration_integration.py for parallel execution
- Update test_remote_executor.py with proper AsyncMock usage

All tests now properly mock async methods and validate parallel execution
behavior when acceleration is enabled.
- Remove 4 obsolete test files (debug logging, subprocess debug, vLLM symlink, redundant HF)
- Add 6 new comprehensive test files covering advanced functionality:
  * test_system_dependencies.json - System package installation
  * test_class_persistence.json - Instance reuse with instance_id
  * test_function_args.json - Serialized arguments/kwargs testing
  * test_mixed_dependencies.json - Combined system + Python dependencies
  * test_class_custom_method.json - Custom method execution
  * test_error_scenarios.json - Error handling and edge cases
- Update CLAUDE.md to fix test file location references

Total test coverage: 11 files (was 5) covering all handler functionality
- Remove custom HfXetDownloader class (~160 lines) - now redundant
- Update huggingface_hub requirement to >=0.32.0 for automatic hf_xet
- Leverage HF Hub's native snapshot_download() with transparent acceleration
- Simplify HuggingFaceAccelerator to use HF's built-in caching and Xet support
- Update workspace_manager to trust HF's cache hierarchy (HF_HOME only)
- Remove manual Xet detection and file-by-file download logic
- Update tests to reflect native HF Hub integration approach
- Add documentation for automatic HF acceleration features

Benefits:
- Automatic chunk-level deduplication via native hf_xet integration
- Simplified codebase with 332 fewer lines of redundant code
- Better performance using HF's battle-tested acceleration
- Future-proof - automatically works with new Xet-enabled repos
- Transparent operation - no code changes needed for acceleration
- Add strategy pattern for HF model downloads with tetra and native implementations
- Implement model pattern matching for selective acceleration
- Add comprehensive test coverage for download strategies
- Integrate with existing workspace and cache management systems
@deanq deanq changed the title feat: Endpoint Persistence using Network Volume and CDR feat: Endpoint Persistence using Network Volume (phase 1) Oct 1, 2025
@deanq
deanq marked this pull request as ready for review October 1, 2025 19:44
@deanq
deanq requested a review from pandyamarut October 1, 2025 19:44
@deanq
deanq requested a review from jhcipar October 1, 2025 19:44
Comment thread src/dependency_installer.py Outdated
Comment thread src/huggingface_cache.py
if repo.repo_id == model_id:
# Check if the specific revision is cached
for rev in repo.revisions:
if rev.commit_hash == revision or revision == "main":

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤔 I think this would short circuit a cachce hit to main even if I had specified a different revision
Right now it doesn't matter I don't think because the only place this is called is w/o specifying a revision, but maybe a TODO so this doesn't get forgotten

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right. It won't matter soon because this entire module will be removed here after #31

@deanq
deanq requested review from Copilot and jhcipar October 6, 2025 02:28

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR implements Phase 1 of endpoint persistence by removing the workspace runtime system from network volumes and simplifying download acceleration. The changes remove workspace-specific persistence (like virtual environments on network volumes) while maintaining HuggingFace model cache-ahead functionality through a simplified direct cache approach.

Key changes:

  • Removed WorkspaceManager and all volume-based workspace isolation
  • Replaced complex download acceleration strategies with simple HuggingFace cache-ahead
  • Updated all executors to work without workspace dependencies

Reviewed Changes

Copilot reviewed 35 out of 51 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
tests/unit/test_workspace_manager.py Complete removal of workspace manager tests
tests/unit/test_remote_executor.py Removed workspace manager integration from executor tests
tests/unit/test_huggingface_cache.py Added new tests for simplified HF cache-ahead functionality
tests/unit/test_hf_download_strategies.py Removed complex download strategy tests
tests/unit/test_function_executor.py Simplified function executor tests without workspace dependencies
tests/unit/test_dependency_installer.py Removed workspace manager dependency from installer tests
tests/unit/test_class_executor.py Simplified class executor tests without workspace dependencies
tests/integration/test_runpod_volume_integration.py Removed volume workspace integration tests
tests/integration/test_python_path_integration.py Removed Python path workspace integration tests
tests/integration/test_hf_strategy_integration.py Removed HF download strategy integration tests
tests/integration/test_handler_integration.py Updated integration tests for new HF cache-ahead system
tests/integration/test_download_acceleration_integration.py Removed complex download acceleration integration tests
src/workspace_manager.py Complete removal of workspace management system
src/test-handler.sh Fixed test file path from test_.json to tests/test_.json
src/remote_executor.py Removed workspace manager, simplified to use HuggingFaceCacheAhead directly
src/logger.py Updated documentation to reflect new namespace
src/huggingface_cache.py New simplified HF cache-ahead implementation
src/huggingface_accelerator.py Removed complex HF acceleration system
src/hf_strategy_factory.py Removed download strategy factory
src/hf_downloader_tetra.py Removed Tetra download strategy
src/hf_downloader_native.py Removed native download strategy
src/hf_download_strategy.py Removed download strategy interface
src/function_executor.py Simplified to remove workspace dependencies
src/download_accelerator.py Removed complex download acceleration system
src/dependency_installer.py Removed workspace manager dependency
src/constants.py Simplified to remove workspace-related constants
src/class_executor.py Simplified to remove workspace dependencies
src/base_executor.py Removed base executor abstraction
docs/System_Python_Runtime_Architecture.md Updated architecture diagrams
docs/Endpoint Persistence.md Added new endpoint persistence documentation
docs/Centralized_Log_Streaming_System.md Updated sequence diagram
Dockerfile-cpu Fixed COPY command for src directory
Dockerfile Added HF_HUB_ENABLE_HF_TRANSFER environment variable and fixed COPY command
CLAUDE.md Updated documentation to reflect new architecture
.github/workflows/ci.yml Streamlined CI workflow

Tip: Customize your code reviews with copilot-instructions.md. Create the file or learn how to get started.

Comment thread src/dependency_installer.py Outdated
Comment thread src/remote_executor.py Outdated
deanq and others added 2 commits October 5, 2025 19:30
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
@deanq
deanq requested a review from Copilot October 6, 2025 02:40

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

Copilot reviewed 35 out of 51 changed files in this pull request and generated no new comments.

Comments suppressed due to low confidence (1)

src/remote_executor.py:258

  • The code after the else block was removed but the same logic is duplicated below. This creates unreachable code as the function returns in the if block above.
        # All tasks succeeded
        return FunctionResponse(
            success=True,
            stdout=f"Parallel installation: {success_count}/{len(results)} tasks completed successfully\n"
            + "\n".join(stdout_parts),
        )

Tip: Customize your code reviews with copilot-instructions.md. Create the file or learn how to get started.

@deanq
deanq requested a review from Copilot October 6, 2025 03:03

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

Copilot reviewed 35 out of 51 changed files in this pull request and generated no new comments.

Comments suppressed due to low confidence (1)

src/remote_executor.py:258

  • Unreachable code detected. The else block at line 252 has been removed, but the code that follows is duplicated logic that will never be reached due to the return statement above it.
        # All tasks succeeded
        return FunctionResponse(
            success=True,
            stdout=f"Parallel installation: {success_count}/{len(results)} tasks completed successfully\n"
            + "\n".join(stdout_parts),
        )

Tip: Customize your code reviews with copilot-instructions.md. Create the file or learn how to get started.

@deanq
deanq merged commit f59bec2 into main Oct 6, 2025
12 checks passed
@deanq
deanq deleted the deanq/ae-1092-tetra-volume-warm-cache branch October 6, 2025 16:02
This was referenced Oct 10, 2025
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants