HN.zip

MXC - a sandboxed code execution system

165 points by nreece - 78 comments
dannyw [3 hidden]5 mins ago
This looks pretty decent actually. Sure, you could consider it a frontend/SDK for bubblewrap/seatbelt/processcontainer; but setting em up consistently is far from trivial; and hand rolling is a really bad idea (speaking from experience).

I like the ‘learning’ mode for figuring out what perms/config a runtime needs, the MIT license, the clear optional telemetry disclosures, and somewhat light and still readable documentation.

Regardless of your views on Microsoft, this looks quite useful; serves a clear purpose, and from a quick glance, looks like a high quality project even if it’s just the first version.

saghm [3 hidden]5 mins ago
I've toyed around with bubblewrap before, and I agree that it becomes pretty unwieldy when you start trying to do anything moderately complex. It's a giant bin of Legos that you can put together however you want, but in practice as an end user you probably just want a few specific sets that come prebuilt most of the time.
LeBit [3 hidden]5 mins ago
It does look nice and easy.

Not sure I will give up on smolvm though.

I have scripts to launch an instance per task. Nono wraps my coding agent.

I provide the git clone for the specific task.

It works really well.

binsquare [3 hidden]5 mins ago
thanks

I'd also add that MXC is still sharing the host kernel.

i.e. on mac MXC uses seatbelt, so the agent shares your kernel.

but smolvm gives each task its own VM with its own kernel, same on mac, linux and windows. + additional layers like preconfigured seccomp, landlock, wherever it makes sense

chrisweekly [3 hidden]5 mins ago
Yeah, smolvm (from https://smolmachines.com) is awesome. Have you written up your setup in any detail?
gandreani [3 hidden]5 mins ago
I agree! I can't comment on the design of the SDKs but from the looks of it this is a great option to integrate into agent harnesses/pipelines.

The "audit" and "debug" mode are specially useful. I use bubblewrap and using a new harness with it is usually a couple of rounds of wack-a-mole with strace to figure out all the harnesses dependencies.

simonw [3 hidden]5 mins ago
This is promising, but there's one feature that's missing that I really care about: fine-grained networking.

They have this for Windows and Linux, but it's sadly missing for macOS - see the support table here: https://github.com/microsoft/mxc/blob/main/docs/backends/sea...

Things macOS is missing include "Allow/deny by hostname" and "Allow/deny by IP, CIDR, port, or protocol".

The rest all looks great, and if you are on Linux or Windows those restrictions don't apply.

I guess this is the universal challenge of building an abstraction layer over multiple different technologies.

gregwebs [3 hidden]5 mins ago
Fine grained network policies is supported by microsandbox- a project that has already been working hard at building an abstraction layer over multiple different technologies. Microsandbox (on unix) builds on top of libkrun (a VM abstraction layer for unix). I am building a convenient runner on top of microsandbox: https://github.com/runcontain/runcontain (undergoing a rename right now). The best thing Microsoft could contribute right now would be great technology for light-weight containment on Windows.
binsquare [3 hidden]5 mins ago
hey simon,

smol machines actually support exactly those things across macs,linux, windows btw: https://github.com/smol-machines/smolvm/blob/main/AGENTS.md#...

here's a snippet of how it looks like to configure that:

  [network]
  allow_hosts         = ["api.github.com"]          # hostname, also allows its subdomains
  allow_host_patterns = ["example.com", "*.npmjs.org"]  # exact names, or *. for subdomains only
  allow_cidrs         = ["10.0.0.0/8", "1.1.1.1"]  # IP ranges  or single IPs

  [[network.credentials]]
  name                 = "github"
  environment_variable = "GITHUB_TOKEN"
Bnjoroge [3 hidden]5 mins ago
+1 on smolvm! I provision a smolvm per job in preloop, a local/self-hosted github actions( https://github.com/preloopdev/preloop). iirc there was also some work to add a deny-cidr option as well which would give you more flexibility. But the egress filter captures most of what you need so it works great nonetheless.
binsquare [3 hidden]5 mins ago
love seeing preloop's work :)
githubnuoo [3 hidden]5 mins ago
how does it do it?

proxy in the middle (but cert pinning problems)

or DNS filtering? (but agent could have "memorized" stable IP)

binsquare [3 hidden]5 mins ago
Cert pinning: not a problem. we don't decrypt traffic, we just pass it through. The only exception is hosts you give a credential to, since we have to insert the key.

Memorized IP: doesn't work, the vm can only connect to an IP if it came from a DNS lookup of an allowed name. Any other IP is blocked.

A bit of "shared responsibility" philosophy kicking through but I try to have good defaults

dannyw [3 hidden]5 mins ago
They could bundle in a HTTP proxy (enforcing similar rules) perhaps. It takes a bit of reading to dig-through the Claude speak, but "Egress confinement is enforced; using the proxy is cooperative" simply means that there's no network egress, except through the proxy.

Of course, that only limits HTTP; and not other forms of network requests.

simonw [3 hidden]5 mins ago
... interestingly, Anthropic's SRT is built on the same macOS primitives and DOES support the network configuration I'm looking for:

https://github.com/anthropics/sandbox-runtime/tree/main#as-a...

  const config: SandboxRuntimeConfig = {
    network: {
      allowedDomains: ['example.com', 'api.github.com'],
      deniedDomains: [],
    },
    filesystem: {
      denyRead: ['~/.ssh'],
      allowWrite: ['.', '/tmp'],
      denyWrite: ['.env'],
    },
  }
amluto [3 hidden]5 mins ago
Sadly, skimming the docs about how the different backends work makes me think that almost this entire project was done by a recent-gen LLM that interpreted its instructions as “make these things work at all costs” instead of “make a considered design that cleanly and securely fits its use case”.

Even the bubblewrap integration docs are basically a stream of consciousness vibe splat. I have approximately zero confidence in the results.

neobrain [3 hidden]5 mins ago
Do any of these sandboxing solutions have a dynamic component to them that lets you grant permissions, starting with a minimal sandbox and asynchronously adding permissions as they become necessary? Harnesses try to do this when accessing non-project folders, but it's not always strictly enforced and generally not revocable. Harnesses also block agent execution until a decision is made, which requires constant monitoring to ensure progress can happen when the agent could easily proceed with an alternative method right away.

I like the idea of a minimal sandbox that protects against accidental `rm -rf` and against personal data leakage, but such a setup then often gets in the way of the specific task to be done. Ideally the sandbox would be able to aggregate blocked accesses and then expose them in an external TUI dashboard, where I can then enable access (without blocking any running agent on this, since that's prone to "press okay" fatigue).

Does anything close to this exist yet?

0kk33 [3 hidden]5 mins ago
If I understand you correctly https://nono.sh/ might go into that direction. It can add permission after the sandboxed command is terminated based on which blocks occurred. Its not life though as you seem to describe
neobrain [3 hidden]5 mins ago
Very interesting concept to make permissions specific to individual shell commands though. Certainly good that people are experimenting with these approaches, hopefully ideas will eventually converge so we don't need to know like a 100 different sandbox projects :)
Neywiny [3 hidden]5 mins ago
I couldn't get cline TUI to run in it though. Apparently node does fancy temp folder things and it just wouldn't go
__MatrixMan__ [3 hidden]5 mins ago
If only we had a decent capability based OS. Then all processes would have the property you're after. If you want it to be able to write to a disk, dependency inject a disk writing capability. Need to talk to a remote host, provide a handle for just that host.

Don't want these capabilities? Do nothing, that's the default state.

pjmlp [3 hidden]5 mins ago
Unfortunately, capabilities based OSes, or written in mostly safe systems programming languages aren't by lack of trying.

However outside mainframes and micro-computers, or niche deployments, adoption has been a challenge.

sleepybrett [3 hidden]5 mins ago
iOS/macos and linux both support capabilities, now the depth of those capabilities in all cases might lack depth but people are moving in this way.
__MatrixMan__ [3 hidden]5 mins ago
Let's take filesystem access on linux for example. You can run as a user without permission, or you can configure something like apparmor/SELinux to stop the program if it attempts do do something not on the list, or you can use containerization to build a limited world for the program to see...

But isn't this all a bit adhoc? The fundamental presumption is of unrestricted capability and then the OS provides options for restriction.

Perhaps I've misunderstood it but I think capabilities are inside out from all of that: You need to walk, so here are some legs. You need to swim so here are some fins. As-is it's more like: you can't go over there so here's an ankle monitor.

pjmlp [3 hidden]5 mins ago
As does Windows, that doesn't mean they are actually used as they should.
shwaj [3 hidden]5 mins ago
Fuchsia is an “option”, although it doesn’t support much hardware without writing your own drivers.
brendank310 [3 hidden]5 mins ago
Not associated with the project but have been following the development, https://5bsd.org/ could be a good answer. Kory’s been adding capability enforcement for the Linux APIs on top of BSD kernel (among other cool features).
ghm2180 [3 hidden]5 mins ago
I think this may be a false choice because of the work pattern you might be used to. Assuming here, If you work in interactive sessions where it's open ended there is a boundary where you have done enough research/prototyping and you need to move to implementation and the permission scope has to change now. I think the realization that you might have is that if you're doing this then it's probably best to separate the automated AFK part from your initial research part.
neobrain [3 hidden]5 mins ago
Even for pure research/prototyping, you quickly run into the problem that your sandbox is either prohibitively minimal or overly permissive. Depending on the exact task, you may need GPU access, Docker/Nix socket access, ability to ptrace processes (gdb), run webfetches, etc. If I define a "research" sandbox profile to allow all of these, I might as well not have any sandbox at all.
ghm2180 [3 hidden]5 mins ago
I have faced this exact dilemma as well. It's not always clear what permissions are needed in advance for my pi sessions and it's child sub sessions. A simple example is when child sessions do a task they locally want to fire random docker commands to learn the state of my local docker devstack.
agentdev001 [3 hidden]5 mins ago
Take a look at Nvidia OpenShell, and their dev blogs on it. That seems like what you're looking for.
neobrain [3 hidden]5 mins ago
Sounds OpenShell locks filesystem access on sandbox creation - arguably the most important isolation feature, at least for my use cases. Architecturally it looks right though!
spankalee [3 hidden]5 mins ago
This really should be a WebAssembly runtime so that the same computations can run on different platforms and devices.
syrusakbary [3 hidden]5 mins ago
We have Wasmer :)

https://wasmer.io/ (some cool examples here https://wasmer.sh/)

spankalee [3 hidden]5 mins ago
Wasmer is one of many WebAssembly runtimes, not sure why it needs specific hyping here.

I tend to use Wasmtime, V8, WAMR, and browsers. The point of WebAssembly is that there are many runtimes.

kernc [3 hidden]5 mins ago
350,000 of mostly Rust SLOC [1] ... And the upstream sandboxes aren't even vendored!

I'd be way more confident building upon something I can grasp and understand. [2]

[1]: https://ghloc.dev/microsoft/mxc [2]: https://github.com/sandbox-utils/sandbox-run

dannyw [3 hidden]5 mins ago
If you look around the files, I think at least half is comments or unit tests, e.g.

https://ghloc.dev/microsoft/mxc?branch=main&locsPath=%5B%22s...

That site thinks this file has 2.9k sloc and doesn't seem to parse rust comments. In reality, there's only 1,465 sloc; and 635 loc of tests.

Definitely nowhere near 350k sloc.

--

As for your sandbox run: it's a single-contributor project, seems to have only have basic smoke tests, and has a few major/critical security issues:

* _generate_seccomp_filter compares newline-deliminated syscalls, against a multi-line blocklist, meaning the entire function doesn't block anything and is essentially a no-op.

* Main script invokes working directory's .env as shellcode, before switching into restricted filesystems and dropping capabilities. Attacker-controlled .env can run shellcode with full privileges.

* Lots of race conditions which I haven't verified, but doesn't really matter.

I'd make PRs, but I don't think it's a good idea to try and DIY a sandboxing system in bash with minimal SLOC as the target in the first place. I'm also slightly concerned that most of your comments on HN seem to be promoting this repo?

kernc [3 hidden]5 mins ago
> at least half is comments or unit tests

Thanks, I see there's a slight (~50%?) overestimation there, but then again, even unit tests and comments in a target programming language count as syntactically correct code that needs to be evaluated and reasoned upon. I'm not that familiar with Rust's runtime introspection features, but in languages like Python, even the comments can directly affect code (e.g. `Foo.__doc__ = Bar.__doc__ + SOME_ANNEX`).

> _generate_seccomp_filter ... the entire function doesn't block anything

Many thanks! I've applied a fix—it's a single line added. The missing test is pending a runnable that invokes one of the forbidden syscalls. As I have no qualms about force-pushing around a repo that nobody forks, happy to credit you(r LLM) proper!

> Main script invokes working directory's .env as shellcode

The sandboxed process can't overwrite existing .env files [1], but it could create a new $PWD/.env file, hoping to "escape" at next sandbox execution. That's a valid concern I'll have to think about some more.

[1]: https://github.com/sandbox-utils/sandbox-run/blob/c97d065184...

> Lots of race conditions

I sometimes experience "Slirp not ready in time" [2], but it's due to a so far unexplained upstream issue [3]. I you have time/tokens to spare, I'd appreciate those PRs and further similar feedback!

[2]: https://github.com/sandbox-utils/sandbox-run/blob/c97d065184... [3]: https://github.com/rootless-containers/slirp4netns/issues/35...

Don't know whether it's a good idea. It sure has got its issues. But even as the SLOC count and the number of bugs metrics are proved correlated in literature [4], min SLOC is not the primary target—a reasonably graspable and stable composition of few dependencies is. Whereas overreliance on third parties nowadays often ends with a rug pull one way or another. We simply can't count on this "MXC" (...) to be maintainable/non-archived even a year from now, just when I'd get it all properly integrated and set up.

[4]: https://softwareengineering.stackexchange.com/questions/1856...

> slightly concerned

Oh, I certainly wouldn't like to limit myself to promoting just this repo! ^D^ HN is a good venue, lots of smart people around! I see everyone shilling their own sh** all the time. Often in green usernames. :shrug:

IshKebab [3 hidden]5 mins ago
And it's 600 lines of dense Bash. I trust 350k lines of Rust way more than that!
kernc [3 hidden]5 mins ago
600 lines of dense POSIX Shell—in some respects that's even worse!

You would not be aware of the amount of trust you are putting into that.

its-summertime [3 hidden]5 mins ago
I feel a better metric is `lines changed / time` as that affects what will be audited as time goes on

That being said, a month of MXC has more line changes than 2-3 years of runc

zmmmmm [3 hidden]5 mins ago
I do hate adding abstraction layers needlessly, but currently this does look like it might solve a real problem for me. Or at least the concept of it.

I want to support sandboxing for my app, and commands it launches, but there is nothing that actually works across all the environments I want code to run in. So yes, if I had one tool that could be an abstract interface and let the user set up and configure their sandboxing completely separate to my app, and do that at run time based on declarative policy - it would be handy.

epage [3 hidden]5 mins ago
Been looking at sandboxing, both low level and higher level like this.

The API for their Rust mxc-sdk looks nice but

- their "sdk" has binaries and the build script has logic for them

- their build scripts do windows-exclusive work on all platforms

- not putting some of the backends behind features causes more build script work (and that work will break on future Cargo versions)

- at least some of the remaíning build script work doesn't need to be a build script

- it seems pretty dependency heavy

pprotas [3 hidden]5 mins ago
Microsoft stole my idea :) (joking obviously, everyone and their mom is making sandboxes) https://github.com/pprotas/slopbox
eminence32 [3 hidden]5 mins ago
A little off topic, maybe, but I've been having great luck with wasmtime and wasm32-wasip3 for writing sandboxed plugins. The tooling is pretty nice when you write plugins in rust, but I don't know what it looks like for other languages right now.

wasip3 is not stable yet, but it has a lot of nice changes (compared to wasip2) for integrating with async code

minraws [3 hidden]5 mins ago
Why is everyone making their own code execution agent runtime engines I have an entire project built on top of openshell already, why not first come up with a sandboxing policy design, like unix did, and then build on top of that.

Currently all project do tend to agree on what and how they work but certain things being different makes porting tedius, if all of them have a bare minimum subset common amongst them it would be much easier to switch, and validate security surface area.

I feel like there are more vulnerabilities in this vibe coded slop sandboxes, and it's more likely everyone one of us trusting them to build projects around them will shoot our foot off once a cve is hit in one that's common in all of them but since they are all slop copies someone will have to figure out how they apply to all others and then manually fix it properly, and if one of them makes a CVE public it will leave dozens of these runtimes open to exploits.

I wish the best to my future self with regards to security I feel like we are completely screwed. Since we can no longer depend on upstream for security.

lifeisloving [3 hidden]5 mins ago
I saw a tweet that said:

"Im really fkin worried we're all building the same thing"

Everyone has been building a harness/sandbox the last 6 months. Ive seen dozens and dozens shared in discords.

Even companies are totally stuck focused on the same paradigms.

The previous iteration of this was RAG/Chat interfaces. See PewDiePie's project. Last month it was briefly everyone building the same classifier.

Peter Thiel, gave a lecture about this same phenomenon 15 years ago likely because he observed the same things going on during other hype cycles. Everyone building the same things. Its called something like "Dont build the obvious thing"

This is why Im moving towards hardware for personal projects, it forces me to be much more creative and think outside the "How can I make something AI adjacent/powered" trap thats so easy to fall into in pure software right now.

nozzlegear [3 hidden]5 mins ago
Harnesses and sandboxes are the new JS framework of the LLM era.
hobofan [3 hidden]5 mins ago
Different use-cases have different requirements.

e.g. this one puts multi-platform support as a high requirement, a requirement that OpenShell doesn't fulfil (and likely won't given it's architecture/goals).

torginus [3 hidden]5 mins ago
The problem with OS level sandboxes, and the reason why WebAssembly's being explored in this space (and Electron is so popular), is that relying on OS/hardware features means your TAM shrinks to a fraction of total, and it's historically well known you set yourself up to lose.

History is littered with tons of super cool OS features that didn't manage to gather enough market share and ended up as cool futures, and fodder for 'we invented the future 20 years ago' style articles.

fg137 [3 hidden]5 mins ago
My guess is that Microsoft thinks this can be deployed with standard, company-wide policy across platforms (mostly) with their IT management tools which poses a unique advantage.

In reality, however, knowing how much difference there is between OSes, how tricky it is to configure these things to make them actually useful, and how bad Microsoft products are, I'm not enthusiastic about this project -- there are so many others on the market already, and I'll wait to see if this gains traction.

(Notice that on MacOS it only supports seatbelt? That's not nearly the same as microvm.)

rock_artist [3 hidden]5 mins ago
That’s exactly it. There should be some permission logic for delegating.

But as always, there are rivals trying to set their tone on what’s the standard. We all wish there was one unified agreed concept that will work but I guess the most common one will eventually survive.

Just as Microsoft in a sense embraces Linux with WSL and also Apple has their virtualization framework.

I hope we’ll eventually get unified model management system to include also permissions designed properly

chneu [3 hidden]5 mins ago
Right you are, Ken!
flufluflufluffy [3 hidden]5 mins ago
Yes! My friend and I used to say that all the time
nizbit [3 hidden]5 mins ago
There it is! :)
neilsimp1 [3 hidden]5 mins ago
What do we always say?
thomas536 [3 hidden]5 mins ago
Don't get eliminated!
sharts [3 hidden]5 mins ago
When do you use this? You’re writing the app so you want the app to sandbox itself instead of… just running in a container?
bri3d [3 hidden]5 mins ago
I think the idea here is cross-platform / runtime pluggability: an app can specify what policies it needs and this thing will map those policies onto the specific runtimes available (containers or VMs). It's basically a container security policy orchestration layer, I suppose.
mintflow [3 hidden]5 mins ago
Seems aws also announced a sandbox solution

I used agent over 1 year and basically always give codex full permission on each thread, do not get issue so far

Why we need this layer of complexity? Or its mainly for big company that need control ?

hedgehog [3 hidden]5 mins ago
It's very useful to ensure that code under test has limited access to resources both to avoid making a mess outside the intended workspace. Otherwise there's risk of deleting or killing stuff it shouldn't, or just using too much memory or CPU and causing OOM kill or other issues. If a runaway command turns into a nice error for the agent then it becomes something that will self-resolve without fuss.
pprotas [3 hidden]5 mins ago
The main usecase for an average developer is preventing confused agents making mistakes like removing sensitive folders, resetting git branches or using API tokens they shouldn't be using
dannyw [3 hidden]5 mins ago
A ~month ago, auto-review (rightfully) blocked a rm that would've nuked my home directory, due to shell mangling (amongst other issues).
joshuanapoli [3 hidden]5 mins ago
If you have a custom agent in a product, then it needs isolation to be sure to protect the customer data.
stingraycharles [3 hidden]5 mins ago
People are running custom agents that are able to run custom code in production just like that ?

I’d personally opt for SELinux in such cases

wild_pointer [3 hidden]5 mins ago
What's also interesting is that they added the Experimental_CreateProcessInSandbox API to Windows, like, last month.

https://learn.microsoft.com/en-us/windows/win32/secauthz/cre...

Funny, now that it's documented, the API name will be stuck with this name forever.

mrpippy [3 hidden]5 mins ago
The docs are pretty clear that it’s experimental, and it’s not even exported from a DLL. It’ll be hilarious if some application depends on it and they have to keep it though
shados [3 hidden]5 mins ago
Is this playing in the same space as the like of gvisor?
arj [3 hidden]5 mins ago
Would this allow a sandboxed container on windows to still run commands in wsl?
plq [3 hidden]5 mins ago
Both Firefox and Chromium have battle-tested cross-platform sandbox implementations, but I imagine it'd be laborious to integrate them as a third-party dependency to other projects. Why not spend resources on repackaging them in an easy-to-use SDK instead of reinventing the wheel?

Microsoft is already a Chromium contributor so it's not like they lack in-house expertise or something.

https://wiki.mozilla.org/Security/Sandbox

https://chromium.googlesource.com/chromium/src/+/HEAD/docs/d...

fassssst [3 hidden]5 mins ago
Codex uses this on Windows now.
superxpro12 [3 hidden]5 mins ago
Right you are, ken!

But seriously, we need docker for models like years ago. I dont want these things running with the ability to run rm -rf /

its should be treated no different than wget | sh

ranger_danger [3 hidden]5 mins ago
How does this compare to Sandboxie?

https://github.com/sandboxie-plus/Sandboxie

rfgplk [3 hidden]5 mins ago
They have a sandbox escape in there. Likewise capability ordering is wrong. Exactly what you should expect from Microsoft.

Since this apparently wraps bubblewrap (another incompetent action on behalf of Microslop), did a quick sweep of that codebase too. Setuid is wrong, capability dropping is wrong, bubblewrap does _not_ protect against compromised/vulnerable kernels (and it should fyi), wrote up a full bubblewrap/mxc sandbox escape too.

zbentley [3 hidden]5 mins ago
Could you link to issue reports to back that up (or maybe file them if these are novel findings)?

If that’s too much of an ask, at least reference the code you found for things like “mxc sandbox escape” or “bubblewrap setuid is wrong”. Those claims require evidence.

smitty1e [3 hidden]5 mins ago
Asked Grok the difference between mxc and flatpak:

"So MXC is a cross-platform “what may this workload touch?” layer aimed at agents. Flatpak is a Linux app format whose sandbox happens to share a backend with MXC on Linux."

zenapollo [3 hidden]5 mins ago
Saw the M stands for Microsoft and immediately closed the tab. 1 it’s unnecessary - communicates nothing but look-at-me branding. 2 toxic company.