Repository navigation
rfc: shelve the Mountpoint fork (not a current priority) - #26
jameskurz-filecoin wants to merge 6 commits into
Conversation
There was a problem hiding this comment.
General feedback:
- Broadly I want to get this stuff going as much as you do, and if the ask is just to get access to the staging region forge to start testing that's fine. We don't want to block.
- There's a lot of AI-ease in this RFC, which makes it very hard to follow -- I can mostly parse because I have existing context on posixmount, but keep in mind others are coming to this blank -- no knowledge of what you've been building and what it does. Given that you're asking folks to take the time to read this, it's important to explain things in simple terms to someone with not a lot of context.
- I've left a bunch of feedback but the key blocking changes or things I at least need a response on before we move forward are:
a. Why do we need to deploy software on the server at all? Unless I'm mistaken this is a client, and I would expect it to work with the regional endpoint without deploying on server.
b. given that the code being deployed is a fork of mountpoint, I'd like to understand why we're not just using mountpoint, since we're taking on 100k lines of Rust code.
|
|
||
| ## TL;DR | ||
|
|
||
| I propose that Fil One stage a deliberately small filesystem-access product: |
There was a problem hiding this comment.
non-blocking feedback: for an RFC like this, I would explain what you've been working on and how you propose to integrate. there's a lot of shared context assumed here -- filfs, core, mention of clockwork that the team doesn't already know about.
the goal here is simple human language other folks can understand.
There was a problem hiding this comment.
Rewrote the top of the RFC as: here is Mountpoint, here is the fork I already made, here is the much smaller thing I am actually asking for. Clockwork is gone from this document.
| object storage has acquired atomic rename, random in-place write, append, | ||
| persistent POSIX metadata, or distributed file locking. | ||
|
|
||
| The first staging workload should reproduce the long-running, multipart, |
There was a problem hiding this comment.
non-blocking reading feedback: what does Fil-281 say in a single sentence? I assume the AI knows, but now the reviewer has to go read track down and this as well.
There was a problem hiding this comment.
FIL-281 in one sentence: a long Proxmox-style backup over a mount died mid-multipart, and we could not finish a restore. Native S3 to the same bucket worked. Now in the Why section.
|
|
||
| ### Direct mount | ||
|
|
||
| `filfs` receives a region, existing bucket, embedded region-profile ID, and an |
There was a problem hiding this comment.
non-blocking feedback: there are a lot of sentences I just don't understand reading this.
"and limits from
the release-bound profile,"
"Production artifacts accept only profiles embedded into the signed release."
"the first
release fails closed instead of claiming it."
"An operator-supplied profile is a development facility behind an explicit
unsafe flag"
It sounds like the binary somehow bakes in the credentials? That merits a simple explanation -- what is filfs, how is it different from mountpoint, and what is it doing?
There was a problem hiding this comment.
Those sentences were about a compile-time config blob (endpoint URL, allowed operations, size limits), not credentials. Credentials are a normal AWS key file. I dropped the jargon. The first test does not use that profile machinery at all.
| | Commercial configuration and workflow | Clockwork after the applicable domain cutover; the existing Fil One path before it | No runtime dependency or commercial model | | ||
| | Stripe and billing effects | The single writer named by the applicable account/domain cutover | No billing calls; client OTLP is operational telemetry only | | ||
|
|
||
| The mount never receives the partner-scoped Service Orchestrator Management API |
There was a problem hiding this comment.
non-blocking: again Clockwork hasn't been really introduced to the process so this is confusing. Service Orchestrator doesn't have an API yet as far as I know.
There was a problem hiding this comment.
Removed. Clockwork has no place in this RFC.
|
|
||
| ## Candidate and evidence | ||
|
|
||
| The reviewed upstream base is |
There was a problem hiding this comment.
non-blocking: This should just say "here's posix mount, and these are the PRs doing important things." Broadly before an intergration I'd have PRs merged.
What are we remediating here?
What is a candidate and what evidence are we producing?
Again, I would just rewrite in terms of what you have, what you want, and what you want to do.
There was a problem hiding this comment.
Replaced that stack with: here is the repo if you are curious, and the four PRs should wait. This RFC is no longer proposing we merge them.
| records zero release evidence and zero certifications, so no region or artifact | ||
| is currently described as qualified. No region is certified. | ||
|
|
||
| GitHub-hosted jobs currently terminate before their first step because of the |
There was a problem hiding this comment.
non-blocking style feedback: why do I care about any of this? It seems like this is just details on what is going on with posixmount ways its ready/not-ready. Typically I would assume "don't deploy to staging when its ready" is a given.
There was a problem hiding this comment.
Cut. Staging-readiness of the fork is irrelevant if the first test is upstream Mountpoint on a laptop.
|
|
||
| ## Staging and rollback | ||
|
|
||
| Deploy the exact signed candidate to Forge staging `eu-central-3` using a bucket |
There was a problem hiding this comment.
blocking feedback:
why are we deploying a mount product into eu-central-3?
unless I've missed something everywhere else, the mount product is a client, that would live on a client machine, and use client credentials generated from the console. Is there a server component to filfs I am missing?
There was a problem hiding this comment.
We should not. That was the wrong picture. filfs / Mountpoint is a client you run on the machine that wants the files. It talks to the existing S3 endpoint. Nothing gets installed on Forge. I need a staging bucket and a key; I do not need a deploy into eu-central-3. CSI would be a cluster component and is out of scope here.
| S3-native workloads. It does not help tools that require file paths, so those | ||
| workloads either need a migration adapter or cannot use Fil One directly. | ||
|
|
||
| ### Lightly configure upstream Mountpoint |
There was a problem hiding this comment.
blocking feedback:
not mentioned here is that filfs is a fork of Mountpoint. this is pretty relevant to rest of the proposal, as we're currently taking on maintaining 100,000 lines of Rust code, a language we're not super familiar with, and it's not clear what's materially different from Mountpoint, except that we have to maintain it and keep it in sync with upstream plus whatever forks we've made. I need at least @alanshaw to weigh in here.
There was a problem hiding this comment.
OK, yeah I did not realise that.
In general @jameskurz-filecoin, engineers want to be responsible for the minimum amount of code possible (which is perhaps contrary to popular belief). Can we at least update the RFC to transparent about what posixmount is, and why it is necessary to fork/re-implement mountpoint?
More importantly, I am missing what this is bringing to the table that could not be covered by some good user facing docs/tutorials on how to use mountpoint with Fil One.
As much as I would love to learn Rust, no one on the engineering team here has expertise in Rust and we'd be relying solely on AI to understand the code properly. I think this is a significant risk and it makes me hesitant to adopt the code base.
To that end, I'm not entirely clear on what the proposal is in this RFC, because it's not "should we build this?" - as it already exists. Is it instead "can engineering adopt this code so that we can sell it as a product?".
If my understanding of the RFC is correct (which I'm not convinced it is), I'd like answered:
"We need to provide users with a file mount binary with baked in credentials because..."
There was a problem hiding this comment.
Agreed, and I should have led with this. filfs is a hard fork of Mountpoint v1.23.0 plus the CSI driver. That is ~100k lines of Rust this team does not own. The rewritten RFC parks the fork. First experiment is upstream Mountpoint (and maybe rclone mount). If that is enough, the product is a docs page. If it fails at something docs cannot fix, I will come back with a short patch list — not 'please adopt the fork.'
There was a problem hiding this comment.
That is the right question and the old RFC hid it. We do not need a binary with baked-in credentials. We need to know whether upstream Mountpoint against a Fil One bucket covers the backup/restore case. If it does, we write docs. If it does not, we talk about a small wrapper. We should not take on the fork so we can sell a mount product — and this RFC no longer asks for that.
| I propose that Fil One stage a deliberately small filesystem-access product: | ||
|
|
||
| - Linux amd64 and arm64; | ||
| - direct `filfs` mounts over an existing Fil One bucket; |
There was a problem hiding this comment.
filfs is my fork of AWS Mountpoint. It is a FUSE client: you run it on Linux, it talks S3, the bucket looks like a folder. The rewritten RFC uses Mountpoint as the first test instead of filfs.
| - the documented `core` operation set: read, list, sequential create, and | ||
| delete when explicitly enabled. | ||
|
|
||
| This is an access method over the existing S3 product, not a new storage, |
There was a problem hiding this comment.
Is there not a product that mounts an S3 bucket as a filesystem already? Is it this - https://docs.aws.amazon.com/AmazonS3/latest/userguide/mountpoint.html
There was a problem hiding this comment.
Yes. That is exactly it: https://docs.aws.amazon.com/AmazonS3/latest/userguide/mountpoint.html — and the rewritten RFC says we should try that first, not our fork.
| and | ||
| - Kubernetes workloads that need a static volume backed by an existing bucket. | ||
|
|
||
| S3 remains the native interface and the honest semantic model. The mount gives |
There was a problem hiding this comment.
honest semantic model
This is AI speak - can you translate?
There was a problem hiding this comment.
It meant: the files are still S3 objects. The mount does not magically become a POSIX disk. Rename, in-place overwrite, and file locks are not there. I just say that now.
|
|
||
| ### Direct mount | ||
|
|
||
| `filfs` receives a region, existing bucket, embedded region-profile ID, and an |
| the release-bound profile, then performs an authenticated S3 readiness probe | ||
| before reporting the mount ready. | ||
|
|
||
| Production artifacts accept only profiles embedded into the signed release. |
There was a problem hiding this comment.
I'm confused, does filfs provide binaries with baked in credentials?
There was a problem hiding this comment.
No. Credentials are a normal AWS key file, same as the CLI. I caused that confusion with 'embedded region profile,' which was config (endpoint + limits), not secrets, and is not part of the first test.
|
|
||
| The mount never receives the partner-scoped Service Orchestrator Management API | ||
| credential and never calls Clockwork. It has no account, order, SKU, price, | ||
| subscription, tenant, or billing state. |
There was a problem hiding this comment.
Right, it's basically just an S3 client with an access key...
There was a problem hiding this comment.
Yes. Client + access key + existing S3 API. No new server.
| ## Candidate and evidence | ||
|
|
||
| The reviewed upstream base is | ||
| [`f708ebd5`](https://github.com/fil-one/posixmount/tree/f708ebd547c90af4fcebcffb434bad34e5cbe38d). |
There was a problem hiding this comment.
Yes. The previous draft made it sound like we had invented a mount product. We did not.
|
|
||
| ## Staging and rollback | ||
|
|
||
| Deploy the exact signed candidate to Forge staging `eu-central-3` using a bucket |
| S3-native workloads. It does not help tools that require file paths, so those | ||
| workloads either need a migration adapter or cannot use Fil One directly. | ||
|
|
||
| ### Lightly configure upstream Mountpoint |
There was a problem hiding this comment.
OK, yeah I did not realise that.
In general @jameskurz-filecoin, engineers want to be responsible for the minimum amount of code possible (which is perhaps contrary to popular belief). Can we at least update the RFC to transparent about what posixmount is, and why it is necessary to fork/re-implement mountpoint?
More importantly, I am missing what this is bringing to the table that could not be covered by some good user facing docs/tutorials on how to use mountpoint with Fil One.
As much as I would love to learn Rust, no one on the engineering team here has expertise in Rust and we'd be relying solely on AI to understand the code properly. I think this is a significant risk and it makes me hesitant to adopt the code base.
To that end, I'm not entirely clear on what the proposal is in this RFC, because it's not "should we build this?" - as it already exists. Is it instead "can engineering adopt this code so that we can sell it as a product?".
If my understanding of the RFC is correct (which I'm not convinced it is), I'd like answered:
"We need to provide users with a file mount binary with baked in credentials because..."
| records zero release evidence and zero certifications, so no region or artifact | ||
| is currently described as qualified. No region is certified. | ||
|
|
||
| GitHub-hosted jobs currently terminate before their first step because of the |
Hannah and Alan could not tell what filfs was, why it was not Mountpoint, whether credentials were baked in, or why anything would be deployed onto Forge. This version answers those directly: it is a client, the first test is upstream Mountpoint, and the ask is a staging bucket plus a key.
|
@hannahhoward @alanshaw thank you — the original draft was hard to read and asked for the wrong thing. I rewrote the RFC from scratch. Short version:
Happy to talk it through live if that is faster than another round of comments. |
Hannah and Alan asked what filfs is and why we would deploy it. After reading the repo: it is a client, credentials are a key file, write-back is off, and the first experiment is stock mount-s3 with path-style to ingot.staging.fil.one. The fork is parked.
Staging is ingot.staging.fil.one, path-style, region eu-central-3.
|
@hannahhoward @alanshaw second pass, after actually reading posixmount. The first experiment is now an exact command: mount-s3 "\$BUCKET" /mnt/filone \
--endpoint-url https://ingot.staging.fil.one \
--region eu-central-3 \
--force-path-styleThat is stock Mountpoint, path-style, from a laptop. |
Not pushed. Off-the-shelf is an acceptable outcome; the fork is not a priority.
|
@hannahhoward @alanshaw please ignore my earlier comments about running a staging Mountpoint test now. I rewrote this RFC again around the actual decision:
The document is the source of truth; the in-thread replies below were from an earlier pass. |
|
@jameskurz-filecoin -- happy to merge but also, I think the right time to merge an RFC is when we're affirmatively adding something. For now, we could simply record this strategy as "the mount approach for now is simply off the shelf S3 Posix adapters, specifically Mountpoint" |
📖 Preview
Folder-mount spike on AWS Mountpoint. Adopting the fork would mean owning ~100k lines of Rust. That is not a current priority. Off-the-shelf is a good outcome.
Ask: do not adopt posixmount. Harvest anything useful, mark it a spike, stop the review. If a customer later needs a folder, try stock Mountpoint from a laptop.