Skip to content

Proposal: @tenki compute backend (Tenki Sandbox microVMs) #3310

Description

@AlvaroDeleglise

Problem / motivation

Metaflow can offload step execution to remote compute via @kubernetes and @batch. I'd like to add support for running steps on Tenki Sandbox — disposable Linux microVMs driven by a small SDK (Sandbox.create()sb.exec("bash", "-c", …)).

For teams already using Tenki, this would let them run Metaflow flows on Tenki microVMs without standing up a Kubernetes cluster or AWS Batch, while keeping the rest of Metaflow (datastore, metadata, cards, @retry, @resources, @catch, resume) unchanged.

Opening this as an issue first per the Core Runtime Contributions section of CONTRIBUTING (the change adds a compute backend and touches metaflow/metaflow_config.py).

Proposed approach

A new @tenki step decorator + backend, modeled 1:1 on the @kubernetes plugin, under metaflow/plugins/tenki/:

  • TenkiDecorator(StepDecorator) — lifecycle hooks, @resources merge (the @batch compute_resource_attributes idiom), datastore validation.
  • A click trampoline (tenki step) + a runner that builds the same bash entrypoint @kubernetes builds (mflog + get_package_commands + bootstrap_commands + save_logs) and runs it via sb.exec(...).
  • A thin, lazily-imported wrapper over the tenki-sandbox SDK — a soft dependency, imported on demand exactly like the Kubernetes/cloud SDKs (not added to install_requires).

The remote-execution contract (the bash string + METAFLOW_* env vars) is backend-agnostic, so most of it is reused verbatim. The Tenki-specific parts:

  • Because a microVM has no ambient cloud identity, the backend forwards the matching datastore credentials into the sandbox as env vars (S3/Azure/GS).
  • Because sb.exec is synchronous, the runner runs it in a background thread and tails structured logs from the datastore on the main thread (preserving live log streaming), then maps CommandResult.exit_code / SDK exceptions to Metaflow's retry semantics.
  • Sandbox cleanup: in-process teardown (finally + atexit) + a tenki list / tenki kill CLI (tagged sandboxes) + a server-side max_duration cap as an eventual backstop.

Packaging

I'm proposing this as a core backend, modeled 1:1 on @kubernetes. It's structured so it could instead be repackaged as a metaflow_extensions plugin (only the two registration lines and the config block would move) if you'd rather keep vendor-specific backends out of core — happy to go that route if you prefer.

The one change to a shared core file is a one-line addition of "tenki" to the hardcoded remote-backend list in metaflow/plugins/pypi/conda_decorator.py(the same list that already contains batch, kubernetes, nvidia, …) so that @pypi/@conda treats @tenki like the other remote backends.

Status (already implemented + validated, for context)

I have a working implementation, validated end-to-end against real Tenki microVMs (with MinIO as an S3 datastore): happy path, plain @tenki steps,@retry, concurrent foreach, @timeout, --metadata=service, @catch, resume, @pypi, and orphan cleanup. 41 unit tests; no regressions to existing tests.

The implementation is ready; per the Core Runtime process I'm filing this issue first and will open the PR next.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions