Markdown Version

Session Date/Time: 23 Jul 2026 14:30

[00:00:05] David Skenazi: Okay.

[00:00:14] Shuping Peng: It's time. So I think it's time for us to start. Okay. Hello, everybody. Welcome to the ATP session. And I'm one of the co chair, Shu Ping, and another co chair, Malorie, is online. So please be aware that this session is being recorded. And first, this is the note well. And if you are especially the newcomer of the ETF, please read this carefully. And because I ETF has some precises and policies, you should follow when you attend the ETF. And there are some guidelines and for the conduct an ETF anti harassment policies and also the IPR. And so if you have any thoughts, and please come to the ADs or the the chairs. We are happy to help. And here is the meeting tapes. And since we are in the middle of the meeting, and just for the in person participants, so when please sign the queue. And and for the remote, please keep your audio or video off unless you are presenting, and please use a headset. And, also, your name and your affiliate. And here's the resources. And, yeah, this is the things. This is our very first session of this working group, and we just we are newly charted. And here's the milestones of our working group. Here, we did update adjustment. So k. So this one price note. Okay. So, basically, I think there is a mistake. But sorry about that. So, basically, we what I did is to we we need to first focus on the repository and the sync synchronization and these two drafts. And then so I just switched a beta to order, and we need to focus on the identifier resolution first and then the URI. So that is what we have seen the comments from the mailing list. And and, also, another thing is I add one more milestones, which is already in the charter current charter. It's about this operational consideration. So that is an informational one. So any comments? Any comments on this? So if there is agreement, and then we will update the milestones. Okay. Okay. Next one is the the drafts. Currently, we don't have the drafts with the ATP tag yet. We have the active drop two active drops. Justin, do you want to comment?

[00:04:16] Justin Richer: Yeah. Sorry. Meet echo wasn't letting me get in the queue a second ago. I would just encourage the chairs to not hold the milestones in their ordering all that preciously. Reality is gonna happen. And so, yeah, the ordering doesn't matter that much.

[00:04:36] Shuping Peng: K. So so we don't you mean we don't make the change? It doesn't

[00:04:46] Justin Richer: Change it if you like them in a different order, but I in in my experience, deliverables happen when they get delivered, not when you plan for them to happen. That's that's all I'm really saying is that this

[00:05:03] Dietrich Ayala: this conveys intention. It does not convey reality.

[00:05:06] Shuping Peng: Okay. That's all. Yeah. Okay. It just based on the current discussions in the meeting is but, yeah, we can leave it. Okay? Here, just to emphasize the UI depends on the identifier. So that is the point. It's the whole point. So, anyway, you are going to so if we start working on the UI, you will reach the identifier problem. So that is the whole point. Okay.

[00:05:47] Justin Richer: I'm sorry. I think you misunderstood the comment that I was making. Put them in whatever your order you want because it doesn't actually matter what order they're in. The working it makes sense for the working groups. Tackle them in different orders. What, Khaliyah? Okay. Yes. There is a natural ordering that can come out of this. And if you're saying we should start tackling this one first because there's that, that's great. Do that.

[00:06:15] Shuping Peng: Okay.

[00:06:15] Justin Richer: And all I was saying is don't hold yourselves to having to finish one before you do another or anything

[00:06:22] Shuping Peng: like that. Sure.

[00:06:23] Ben Goering: Because

[00:06:25] Dietrich Ayala: Yeah.

[00:06:25] Justin Richer: Things land when they land.

[00:06:26] Shuping Peng: Okay. Alright.

[00:06:31] Roman Danyliw: So I'm gonna speak as AD. The we do want, at some point, dates on here. So we'll have to do that today. But at some point, we're gonna we're gonna wanna put some dates on the milestone. I'm looking in the data tracker. There's no dates on here. The one and I'm I'm attempting to quickly read the charter as well because I don't I don't remember it. I don't know if it doesn't sound like the operational considerations document is a requirement. It's just it says deliverables for this working group. So if that's something that at some point the working group says they don't wanna do, then we just need to own up to that. So

[00:07:12] Justin Richer: I just wanted to note last time I had a discussion with, somebody about this, the data tracker was looking to drop dates for milestones because they have proven to be unrealistic.

[00:07:25] Roman Danyliw: Yeah. That's and we're gonna we're gonna try to dates are a management tool. And so as the as a manager, I'm gonna wanna see dates at some point. Okay.

[00:07:37] Ben Goering: I was just told that m can't make the mic in time.

[00:07:39] Shuping Peng: I Emilio. I'm I'm Leah. So he's next. I

[00:07:46] Ben Goering: was told or someone is saying in that in the chat, m is saying that they can't make the queue spot right now, so I should go.

[00:07:53] Amelia: No. I'm here. There you go. Just I was just in the next room, and I had audio on remote. So as far as ordering here, I think the ordering maybe only really impacts with regards to discovering dependencies between different things. So I think there might be a dependency from repository on URI and identifiers as well as a dependency on sync from repository. In which case, I mean, you can't really do sync without knowing what the data is in the repository, so you need that. You need those things first.

[00:08:32] Shuping Peng: Yes. Yes. So I

[00:08:34] Amelia: So maybe

[00:08:35] Shuping Peng: the last one is version. I don't know why. And the the first two, I didn't change. The repository and the synchronization, they are still the first and second. And I only switched a bit, the identifier, the UI. Sorry about this.

[00:08:52] Amelia: Oh, okay.

[00:08:53] Shuping Peng: They did it before.

[00:08:54] Amelia: I I guess I'm reading this milestones slide

[00:08:56] Shuping Peng: wrong. Sorry

[00:08:58] Amelia: about that.

[00:08:59] Shuping Peng: No. No. No. It's my I don't know why I updated before. Ben Go, please.

[00:09:04] Ben Goering: Hey, Ben Go. I had two things. First is, like, I don't act I had I don't understand the proposal because I have a hard time parsing this slide and what it means to deliver one and two first because, like, several things are named two have two next to them. It sounds like you're saying repository first then synchronization. Is that the proposal?

[00:09:22] Shuping Peng: No. No. No. The the repository and the synchronization, they are now changed. That's still the first and second.

[00:09:28] Ben Goering: Okay. The other thing is there's, like, some mention of an architecture doc, I think, on the next slide.

[00:09:34] Roman Danyliw: Yeah. I guess I just

[00:09:37] Ben Goering: wanna point out isn't in the data tracker or the agenda for this meeting. But if it were, I would have reviewed it.

[00:09:45] Shuping Peng: Yeah. So the currently, we you can see two documents on the data tracker. And I tried to add this architecture one, but it it didn't show. I don't know why. Probably, it is because it was expired. Cool. Yeah. And

[00:10:02] Ben Goering: would would make if we're having to decide on an order, which Justin doesn't seem to be true, then the it would make sense to me to start with architecture so that all the other things can even be contextualized. But if Justin's right, the order doesn't matter. And the main thing is, like, if there's an architecture doc, let's get it on the list and in the agenda and and the data tracker. Is it early before this meeting is possible?

[00:10:25] Shuping Peng: Yeah. Yeah. I mean, the the architecture one, I just talked with the authors, and they were saying probably if it's up to the group. If we want it to be active, we can have it there, but they would not to progress it or just to make it informational if you think it is useful. Yeah? Cool. Okay. So this goes to we don't have ATP tag tag, the the drafts with ATP tag yet. So here is also the question to the group, whether we want to have the existing drafts to directly change their title or you want to construct a new compose new drafts? So what do you prefer? Any comments? Yeah. Bango?

[00:11:32] Ben Goering: This is Ben Go. I mean, yeah, a two p tag makes sense, especially given the conversation on the list about names in recent weeks and what these things are called and whether we should rename the protocol. So I think ATP makes sense.

[00:11:46] Shuping Peng: Okay. Daniel?

[00:11:49] Daniel Holmgren: Yeah. I also think it makes to get the ATP tag on it. I would be inclined to just use the existing drafts and bring them into the working group with the ATP tag on it.

[00:11:58] Shuping Peng: Okay. Thank you. So any other comments? Justin, please.

[00:12:06] Justin Richer: So I may have missed. Has there been an official call for adoption of these three documents as working group items?

[00:12:12] Shuping Peng: Not call for adoption for now. It's just the individual drafts. We need to have some documents to start with. And, currently, these individual drops, they are individual drops, but today, you see the tag is not ATP. It's AT. So it can

[00:12:28] Justin Richer: That doesn't matter. It can be goldfish. It only matters when it's a working group document.

[00:12:36] Shuping Peng: Okay. Oh, okay.

[00:12:37] Justin Richer: Yeah. You can just put them in your in your related

[00:12:41] Ben Goering: list.

[00:12:42] Justin Richer: And when you do the call for adoption, that's when you make sure that it maps the things so the data tracker puts it in the right place automatically on

[00:12:48] Shuping Peng: the platform. Now I

[00:12:49] Justin Richer: It's it's convention, not restriction. Okay. Okay. You may be covering this in

[00:12:54] David Skenazi: a bit, but is there going to be a

[00:12:55] Justin Richer: call for adoption for these soon?

[00:13:02] Shuping Peng: I mean the call for adoption for the to become a working group draft? Yes. That depends on the discuss discussions in the That

[00:13:11] Justin Richer: is is fair. Thank you.

[00:13:13] Shuping Peng: Yeah. Okay. So so for now, we leave the drafts as it is. Right? And then we progress them. And then when it reach the rough consensus, we move it to the working group. Right?

[00:13:29] Justin Richer: Traditionally, the rough consensus happens when they're working group documents unless there's the call for adoption is to say, is this a good enough starting point for us to start getting mad about it?

[00:13:41] Shuping Peng: Yeah. Officially. I mean, so currently, we just leave the titles of the drafts as it as it is.

[00:13:46] Justin Richer: Yeah. And Yeah. What I'm saying is that if the if the working group wants to work on these, they really should be working group documents.

[00:13:55] Shuping Peng: Oh.

[00:13:56] Justin Richer: So I I'm I'm confused about what the proposed process here because it sounds like you want to finalize finish the documents, then bring them into the working group. If that's the case, why is the working group here?

[00:14:07] Shuping Peng: I think there might be some clarify?

[00:14:09] Eric Rescorla: Yeah.

[00:14:09] Justin Richer: Yeah. I'm I'm confused. I'm sorry.

[00:14:11] Mallory Knodel: It can be very helpful to have the documents that are associated with the working group show up in the data tracker before they've been adopted so that people like yourselves can go there and see what's being discussed even if the draft is not yet adopted. So that's the use of the tag in the data tracker. So it's perfectly reasonable to say, we want these documents to have the ATP tag so they show up and you know it's being discussed and then adoption being a separate step. So thank you for giving us the opportunity to clarify the difference.

[00:14:40] Justin Richer: Great. Okay. Chairs can just go add those drafts. You they don't need to be renamed.

[00:14:44] Shuping Peng: So yeah. They're already in there. They just Or in there.

[00:14:48] Ben Goering: Is there a?

[00:14:50] Shuping Peng: It's it's more

[00:14:51] Mallory Knodel: of an automated thing. Right? So but, yeah, again, thanks.

[00:14:55] Shuping Peng: Okay. So just one comment is for the new new drafts, please use the ATP tag. Otherwise, it's I have to do it manually. I it won't be possible for me to know you have the you have submitted the draft. Yeah. Okay. That's it. Yeah. So here is our agenda, and we have got six items. So agenda bashing, any changes? No? Okay. So now we can start. Thank you.

[00:15:50] Brian Newbold: Cool. Hello, everyone. I'm Brian Newbold. I work at Blue Sky. I'm currently in the middle of a summer sabbatical for another month or so. I've been working on some of these documents, with my coauthor, Daniel. And so this is kind of a this is gonna be an update of these mostly these two repository and synchronization documents and their kinda current status and how drafting is coming along. And I'm mostly gonna cover I'm not really gonna get too into the the content or changes that might be made. Daniel's gonna give another talk right after this that'll get into this. This is mostly just about scope and, like, did we get everything in here in the charter that we need to have in? Are there things in these documents that maybe should move to other documents? Like, is the structure of this make sense? Should we merge these documents? And then also, like, just what's are there open problems here? Are there open questions that are not resolved? Could you go to the next? Or

[00:16:37] Shuping Peng: No. It doesn't work. It doesn't work. I'll try.

[00:16:40] Brian Newbold: Great. So I'm gonna just give a really quick summary of the charter as I see it. The status of what these two documents are that we have as IDs, and then get into these kinda, like, scope and structure questions a little bit. And then if there's any open issues, it's like a little bit of a call for, like, we could use some help or, like, here's what here's what at at least I feel is, like, unresolved about these documents right now. Yeah.

[00:17:06] David Skenazi: Am I putting it right?

[00:17:07] Shuping Peng: I'm having you.

[00:17:08] Phil Feigl: Okay. Next.

[00:17:09] Shuping Peng: Yeah. It's alright.

[00:17:13] Brian Newbold: Could you do the

[00:17:14] Eli Mallon: next thing?

[00:17:14] Shuping Peng: Okay.

[00:17:16] Brian Newbold: Okay. So this is bringing back the, the bricks slide from the BOF, which is kind of like the landscape of the whole app protocol, framework, I think, is what we ended up calling it at some point. So there's a bunch of components. These gray components are ITF standards that we are building on top of that already exist. They've already been standardized. The green blocks are things that are under like, definitely under the current charter. The yellow things are things that maybe you will recharter in the future and bring in, I'd hope, but aren't work that are in the current scope. And then I added this purple one, which is a thing. I'm gonna come back to that question. It's like, that's not really the current charter. It is kind of an important part of interoperability. It's also a pretty simple piece of the protocol, and maybe we could just pile that into the repository document right now. It working? Okay. It's

[00:18:06] Shuping Peng: you. Yeah.

[00:18:07] Roman Danyliw: We just haven't turned

[00:18:07] Brian Newbold: it on. Okay. Great. So the status the the approach we took was to first start with these documents, basically just taking because we have these written specifications that are just on this random website, appproto.com. So we kind of took those existing specs, rewrote them in the ITF, you know, syntax and format and language and norms. But these are kind of like what's deployed. So the current we wanted to first get drafts that are in the state of how the protocol is running in the network, and then we can kind of, as a working group, discuss if we wanna make any little changes as part of the ITF process. If we wanna get feedback on how things are currently deployed, we can do that as a step on top of that, and there'll be a clear diff. Right? So if we make any changes for the ITF version, it'll look like a little diff. That's just how we approach it. I don't know if that's how we needed to do it. And I think I won't go back over this because we already kinda started it, but we have these three other documents, the URI document, the account identifier system document, and the operational considerations. I think, at least in my mind, these repo and sync are, like, coming along pretty well. They're we're currently really drafting these. We can get these pretty far along. I don't know if we're gonna push it all the way through, but they're feeling like they're getting close to a maybe a call for adoption. And I think next, the URI and account identifier systems, we can start drafting and pushing on those kind of as another phase. And I don't know if we'll get some all the way through to the end before we start pushing on the other ones or we'll end up with a couple of them. There's this untangling these references between the documents. We'll need a little bit of work. These are the current this is just like a little snapshot. I copy pasted out what the outline of the two documents are to kinda what's in each. And I kinda wanna draw attention. There's some of these little appendix things. So these are these other little bits that aren't really core to the repository, but they're things that we're referencing, like which cryptographic curves we're using, these other little identifiers that come along, and some of the details of the JSON CBOR data modeling, like back mapping back and forth between JSON and CBOR. So those are, I think, attached pretty closely to the repository document, that's why we ended up putting those in there. Oops. Skip one. Okay. So there's a couple I'm kind of some of this is just kinda like confessing decisions I made in this drafting and and wanted to get some feedback on if I'm doing this right. This is what these are some of the bigger RFC documents I've ever drafted. So we have this question of which cryptographic systems do we use in App Proto. The current system uses two curve types. Those were just kind of somewhat arbitrary I mean, not arbitrary, but we kinda made these calls, and we kinda wanna not like, is it appropriate to put those in RFCs? Probably not. Where should that decision get made? Where should those live? I put them in the repository document right now. I think the part of the mechanism for making a signature, which is that you take the commit objects, put them in C board format, hash them with SHA two fifty six, and then sign that, that feels like that probably stays in that document. But which of these curves and cryptographic systems should we use, and what's the change process for that? How do we add new curves? I'm not quite sure where that should live maybe in the operational considerations document. That would be one of the later things that we get to. We probably wanna get input from other people at the ITF around this. Daniel's gonna touch on that a little bit more. We've got this data model section as an appendix in the repository thing. Again, I think that this is probably the place to do it. There's this question of floating point numbers, if we should add that to the data model. That'll be another talk later in this session. And then there's some other limits that I don't get into there. We'll probably put in the operational considerations around what's the largest how many bytes can a record be? Gonna be a petabyte? Probably not. But that kind of this is like could use some ITFE feedback on, like, where should those limits? It feels kinda like best current practice, not something we wanna put in a a standards track document. This referencing how to reference between the repo document and the account identifier system, the way that you validate the signatures on a repository is to resolve the cryptographic keys currently associated with the account and use those. I'm not sure if we can just kind of say, this is out of con you know, this is out of scope for the repo document. And then later in that account ID document, describe how that works, or if we need a normative reference between the two and how that how publication goes with that. I just don't have that much experience with that. But I would kinda lean towards just not getting into it too much in the repository document for now and put that reference back in later. There's this blobs thing. I think output of folks or people that have worked with the protocol probably are more familiar with this and know that it's, a pretty small piece of the protocol, but you can have these other media files that are stored on your home server alongside the repository. I think we could specify this pretty quickly. It would just be a small section to add to the repository part. It doesn't feel like it's worth having a whole other document about that later down the line. We could also jam it in the operational considerations, but I'd probably lean towards putting it in in the repo document. So I'm gonna grab water. We have this little thing to untangle. This is another part of the untangling this bricks that there's this framework of everything kind of references on each other, and we we can't we can't standardize everything at once. Part of the synchronization scheme involves these HP paths as things are written right now. We have this XRPC acronym in there, and that's one of those bricks that we might come back to in the future. And so in my mind, that that's like a whole little framework. It's a whole little protocol that I'd rather just not get into right now as part of the synchronization and repository documents. But for interoperability, it's important to know which path to actually call these at. If you wanna download the repository CAR file, you need to call one of these HTTP endpoints. So, know, how do we how do we get something out of this initial charter that's actually usable? People can actually interoperate on, but we're not trying to pull in this new future context. I would lean towards probably putting this in the operational considerations if we can do that, which, like, just kinda leaves the door open to iterating on this in the future. There's also this question of this has this com.appprotonamespace, which is a domain name currently owned and registered by Blue Sky. Maybe we wanna move that to org.itf or some other namespace that's governed in a different way, but that's a whole you know, that's a new I n a registry maybe or some other responsibility that'll need need to be taken on. I would personally lean towards just having something. So the operational considerations feels like a way to put something in that's like, hey. Here's how you do this today. This part might evolve or change over time. You're probably gonna need some legacy compatibility for that going forward, but it would be a way to just kind of move forward. Emilia, do wanna jump in?

[00:25:15] Amelia: Yeah. For the com appproto or appproto.com domain, wasn't that discussed in the context of the trademark for App Proto and bundling the two things together as a this is a thing that is owned by who whoever is overseeing the ad ATP group, whether that was the community, whether that was the whatever the IETF trademark thing is, etcetera, in which case we don't necessarily need to change the the NSID for that.

[00:25:50] Brian Newbold: That's possible. I mean, I think that's just still kind of open work whether where how that will land. But it does you're it's worth noting as you have that that it touches on that discussion. I guess the last thing I would say on that API endpoint question, like, there is this whole schema system and way of kind of, like, defining new API endpoints. I think specifically for these things like get repo that are, like, core to how the protocol operates, at least so far in the kind of, protocol like, development process, we've kinda said, like, hey. These comm dot app proto things, they are special, and the real authority for those is the written specifications. Like, the authority for those will end up being the ITF RFCs, not the resolution process. Like, there's this way to resolve it maybe similar to if you're familiar with old XML schema resolution. And so there's a question of like, oh, do you have to is it the what resolves when you resolve the schema that's authoritative, or is it the specification? I think we want it to be the specification that's authoritative for those schemas in particular. So that may be a way to kind of get around this a little bit for those specific endpoints. As kind of an omission or, like, a way when we were drafting, I tried to not get too much into the PDS and relay concept. In my mind, those are kind of operational. That's, like, kind of the current architecture of the protocol involves those. People talk about them. Developers talk about them a lot, and it might be a missing thing to not define those terms or get into them in the system. But I think we can also potentially do it without having to define those roles. Maybe those roles won't exist in the future, but I've kind of left those out. But that might be a thing other people would object to. And then there's this question of, like, where should we put privacy considerations? I think this was I can't remember, but it might be even mentioned in the charter that we should get into privacy considerations and where should that live. I would propose putting it probably attaching it to the synchronization, like the network wire protocol and things like that. Ted, did you wanna jump in?

[00:27:59] Ted Hardie: Ted Hardy. No relevant affiliation. For the privacy considerations, I would actually suggest that if you're going to have an architecture document, they belong there. Because the overall system will have privacy that or or privacy risks, are not relevant, to one specific aspect, like a particular relay protocol. So I think if you're going to have an architecture document, then, putting it in that is the is the sensible thing to do. If if you're not going to have it, it's probably gonna have to be a section in each of the others. And I strongly suggest you separate it out from the security considerations because it's our our past history that if you try and combine the two, you tend to do a good job at one and not the other. And that is suboptimal given the the remit of the group. And since I bothered to come up to bother people at the mic line, I will go back up to your HTTP endpoints comment and say, I don't think you should change this until there's a noninteroperable change. Right? So keep keep it where it is for now. Once you've made a change to how GET repo works that would not be backward compatible, that would be the moment at which you you would shift it into a new namespace and decide whether that new namespace needs to live somewhere different because you actually want it to be reference to, say, an XML schema maintained at IANA rather than what you would download. IANA actually has a set of that that can be referenced for ITF protocol parameters, and we can talk about that with the folks at INA if you get to that point.

[00:29:37] Brian Newbold: Cool. Thanks. To go back to the architecture document, just to clarify on that, we're not I'm not planning to publish that document. That was only ever just an Internet draft informationally to help the ITF community understand the protocol. If people really wanna push for it, like, the chairs or the director really wanna push for that or they think it's an important thing, I'm willing to maintain it. But right now, it's just a it's just a text file out there on the Internet.

[00:30:02] Ted Hardie: Ted Hardy. I I actually think you you want the information somewhere. If you want to put it as a preference to some other document, I don't really care whether, you know, it it belongs in in some other document or or a separate document. It's easier to rev if if they're separate documents because different pieces actually change at different times, which is actually something that Andy and I were talking about in the in the chat. The same is true for when you're talking about, like, what curves you use for signature. If you keep those in a separate document, if those change, you only have to rev that sec separate document, and that can be a document in which you talk about things like how cryptographic agility is gonna work in the overall system. So I know it seems daunting to have a lot of little documents and to construct ways of of the path dependencies among them. And for the time you're actually writing the first set of them, it's a giant pain. But it turns out for maintainability across the long term, it can be a real advantage if they tend to have changes that are not synchronous changes. If they are synchronous changes, then why bother? But if they tend to evolve in separate paces, that's useful.

[00:31:11] Shuping Peng: Okay.

[00:31:12] Brian Newbold: Thanks. Emilia, do you wanna jump in on this before I

[00:31:19] Amelia: wrap up? Yeah. So for the privacy considerations, I agree with with the previous point that was made about security considerations and merging them. From my experience working on other documents, we tend to have two sections, one for privacy and one for security. And then there are also security BCPs or, best current practice documents that we can publish as well that cover the current state of art. So we can use those as well, I think.

[00:31:50] Brian Newbold: Great. Thanks. I think the queue's locked, but I'm almost done if you can wait.

[00:31:55] Shuping Peng: Just a minute.

[00:31:55] Roman Danyliw: No. No. No. No. You're in there. Sorry. Sorry.

[00:31:59] Phillip Hallam-Baker: Yeah. Yes. Andy asked me to bring this to the mic. Yeah. Really, the endpoint should be a well known service, and we need to get this to tie into SRV. And there's a whole load of things that yeah. I'll I'll I'll make deep more detailed set of comments later, but it's basically just bashing it into alignment with the standards process. But looking through the crypto piece, you know, using these curves and you've got this thing saying, malleability, whatever. One of the reasons that we went through the whole CFRG thing on two curve two five five one nine, It's no better or worse than any other elliptic curve. They're all the damn same. But we did go through all these little nits and hammered them down, and so you can ride on top of that rather than having to hammer them down again?

[00:32:52] Brian Newbold: I like, a thousand percent degree, we do just have this, you know, maintaining current compatibility bots deployed for this protocol. So that's just the balance. Like, if we could start fresh, we'd totally just reference existing work. Okay. So this is, like, a last a little bit of a just burying the lead. I think merging the sync and repo docs would probably make sense. And so unless there's objection, I'm just gonna copy a lot of the stuff from the sync document and put it in the repo document. They awkwardly refer to each other. And I just even when I'm drafting, I have to flip back and forth between the two documents all the time. And I think just having them as one, it'll be a little bit long. But if I look around, there's a couple other RFCs that long. And then, I don't know, maybe we might end up doing something similar with the URI scheme. That seems to be a common pattern. It's just define the URI in line. But at least for now, to start, I think, keep the URI a separate little document. I guess, referencing back to Ted's point, like, that might be something that changes more fast, more quickly, and having it be a separate document could move things along easily. This is a separate little bit of just kinda like open issues. This is a little bit of a call for participation or just kind of issue tracking. There's this one little piece of the MST that we've never really nailed down, which is to define exactly which extra blocks are needed on the Firehose. We kinda do this tautologically right now. We say, include the extra blocks that you need to verify things, but what how you know which blocks you need to verify it is a little open ended. So, we'll see if we can get some more input on this. We've we've taken a whack at it, I think, internally in the past and never really came up with a satisfying definition. So, hopefully, we can come up with a good definition for this to put in the document. And that's it. So I have a little bit of a reminder of a couple of those main points. Otherwise, thanks for everyone's time.

[00:34:43] Shuping Peng: Okay. Thank you. Comment? One comment? No? Okay. Thank you.

[00:34:51] Roman Danyliw: Okay. Thanks.

[00:35:22] Shuping Peng: All the slots for the request has been already taken. Oh, okay.

[00:35:43] Phillip Hallam-Baker: This is working.

[00:35:44] Shuping Peng: Yeah. I Should be. Gave it. Okay.

[00:35:48] Daniel Holmgren: Hey, everybody. Daniel Holmgren, also from the Blue Sky team, and I'm one of the authors on both the repository and synchronization drafts. I'm gonna be going over some of the possible changes for the IETF version of these drafts. These are changes that, like, kind of Brian and I have talked about making to the drafts. I floated these on the mailing list maybe two weeks ago and got some feedback on, them. Some of that feedback has been incorporated into these slides. And kind of the goal of this is to also just talk about, like, the difficulty of making certain backwards incompatible changes in the network, like rating the different levels of difficulty and kind of giving a sense for maybe, like, the scale of the sort of backwards incompatible changes that we're looking at and get a sense of if that is seems too adventurous for the network or if people have a sense that we're going to be making even deeper cuts, to the protocol than that. So a quick overview of that is, kind of just giving a metric for the difficulty of these changes and then getting into some of the changes that I've floated around the repository commit, the cryptographic curves that we support, and the signing scheme, and then a few different things having to do with the synchronization protocol messages framing of the, like, WebSocket subprotocol, and then some things around the cursors for the synchronization protocol. So this is, like, kind of a, like, vague metric of how hard and disruptive certain changes in the network are to make. So AT is, a stateful data protocol. A lot of people are hosting data that is addressed. It's stateful, and we have these, like, strong hash based referential integrity links between them. So that means that changing historical data is often very difficult to coordinate because we have to coordinate this across thousands or tens of thousands of different services, and it also breaks these references. And those broken references have cascading effects because if you try to fix the broken reference and something re like, references the record with the broken reference, then that reference is broken as well. So that's just like a a sense of these things. Things that, like, change the hash of records in the network are the hardest things to deploy. Things that don't touch the hash of records are much easier to deploy. So commit objects are probably the commit of the repository is probably the simplest thing for us to change. That doesn't mean it's not hard to do. We have to roll this out to thousands and tens of thousands of different services in the network. But, ultimately, like, I think that we can do that, and we can get the network to upgrade into that. Kind of similar with the SyncWire protocol. That's a little bit more complex. There's a lot of tooling built around that, but I again, I think that we can coordinate around that. Starting to get into the repository structure in the MST, that starts to get into, like, pretty major data migrations that thousands of developers are going to have to perform over their data. It's not intractable, but, it is much more difficult, than the commit or the sync wire protocol. And then things like the URI format and the record data model, these things are in billions of records on the network right now. I just don't think that we're ever going to be able to fix that. That doesn't mean that we can't touch things like that, but I think that we are going to have to maintain like, I think we are just going to have to specify that there are records out there like that because I think it's just intractable to get those, fixed up. So, now I'm gonna get into a few of the things that, like, that we've chatted about as possible changes, to the to the RFCs. The the RFCs, they're drafted right now, describe the system as it's currently deployed. These are kind of things that we would be looking at for the IETF version. So the this is, like, an example of the repository commit. On the left is how it looks today. On the right is what it could possibly look like to call out the things that changed is we're adding this, like, dollar sign type with the type of object that it is, just providing domain separation on the commit for the signatures, updating the version. Currently, they're version three commits. It would be updating it to version four. We kind of have this, like, did shaped hole in the RFCs that we're drafting. We talk about them in the RFCs as resolvable identifiers currently in the repository commit. They're referenced as DIDs if we want to, like, pull that terminology out of there. Someone on the mailing list proposed using repo for that. Other options could be account, ID conflicts with a lot of maybe, like, programming language, magical words, maybe identifier if we wanna be verbose with it, but but some sort of, term in there. The best I've heard so far is repo, I think. There's also a pre field that was mainly kept around for backwards compatibility with version two. In nearly every commit in the network right now, it's set to null. And I think in this version, we could just formally deprecate that field and say, you don't have to have that in there anymore.

[00:40:39] Shuping Peng: We got people in the queue. Do you want to?

[00:40:41] Daniel Holmgren: Oh, yeah. Sure. Alistair, do you wanna hop in?

[00:40:45] Phillip Hallam-Baker: Yeah. Alistair Woodman. So do you have a plan at all for what happens if things start to be able to crack your algorithms? I mean so The signing algorithm?

[00:40:57] Shuping Peng: Yeah.

[00:40:58] Daniel Holmgren: Yes. I'm actually gonna get that into that on the next slide. Or or I don't know that we have a formal plan yet, but we have we have concepts of a plan. Yeah. So so right now, there's two curves that are supported in the network, NIST two p NIST p two fifty six or sec p two fifty six r one. It's a common choice for EZDSA. It's supported by most HSMs. It seems like a common choice in a lot of IETF protocols. We also support SECP two fifty six k one. The reason for that is that we were looking at, like, kind of the crypto ecosystem and seeing that there's a lot of consumer key management software that supports that curve type in particular. In hindsight, I'm not sure that we really reaped a lot of benefits from that. I think Eric Pescorla pointed out on the mailing list possibly deprecating this. I don't feel, like, super super tied to this. I I think if we were to deprecate it, we would wanna add in another option, at least one other option. And then possibly to add Edwards two five five one nine, there's a lot of developer interest in this. I'm quite interested in doing this. It's a common ask. It's just, like, a lot of work to do, and so we haven't done it. Maybe it's not worth the hassle with quantum stuff looming on the horizon. Maybe we should just if we're, like, looking into upgrading algorithms, maybe we should be putting our energy into that. And then post quantum signatures, I'm not really an expert in that domain. My understanding is that, like, the current IETF recommendation is MLDSA. Yeah. It looks like a couple people in the queue.

[00:42:33] Shuping Peng: Eric. Eric.

[00:42:37] Eric Rescorla: Yeah. I mean I mean, MLDSA is what there is. It's an ITF. So, I mean, now you don't have do SLHDSA. That's for sure. I mean, I can't see the value of adding e d $2.05 $5.00 9 now. That seems like a weird a weird choice. As I sort of said, like, on West, I mean, I think p t I think, sec p 56 k 1 is like a I know it's like the Bitcoin choice, but it's just like a everyone else thinks it's really idiosyncratic, like, choice. So I think, like, trying to keep that trying to get rid of that, like, relatively you know, at some level of of adjustment we've quickly I I don't I'm not sure I'd see why you'd I I think my inclination would be just to, like, add MLDSA and, like, and say we're gonna transition away from sec two fifty six k one and how that not do much else.

[00:43:24] Daniel Holmgren: There are more folks in the queue, or should I keep going?

[00:43:26] Roman Danyliw: There is more.

[00:43:27] Daniel Holmgren: Alright.

[00:43:28] Ori Steele: Hi. I was kinda hoping I'd get a chance to argue with Ecker about signatures, but he says that maybe d two five five point nine isn't worth it at this stage. I that's what I was gonna get in the line to say, basically. Like, and for sec k two fifty six curve, like, the that thing in the upper upper and lower s stuff, like, with hash based structures, like,

[00:43:54] Dietrich Ayala: yeah,

[00:43:55] Ori Steele: I I I really do think it it would be great to get off of it soon. So if MLDSA is too hard, then you could buy the e d two five five one nine argument. It's, you know, very stable in that way. But I think at this point, it it it's better to just go to MLDSA. So I'd be interested in which variant of MLDSA and, like, all all those other details and making sure that we don't have something like the upper and lower s problem with MLDSA, which I'm pretty sure doesn't exist, but I'm not really a cryptographer. Thanks. Great. Oh, and I'm Ori Steele.

[00:44:30] Shuping Peng: Here you go. You want

[00:44:32] Ben Goering: yeah. Maybe. Hey, Megan.

[00:44:34] Shuping Peng: Got two more. So, Pango, please.

[00:44:37] Ben Goering: It seems like there's a choice of, like, is are we gonna fix, like, one or two curves in the repo document, which may be merged with other ones? And, like, there's a choice to make of having an open registry or, like, factoring out the specific curve choice altogether. But if the docs are gonna merge, we can it's, like, hard to even have that conversation right now. So it just feels like we need to either not merge those documents so that but it we need to figure out what the documents even are, adopt them into the group. And as soon as they are in the group, maybe ask for advice from either CFRG or a secure like, some other groups that are familiar with how to do this. Do and I guess the question is, do do you get the sense that you wanna just choose two curves, like one elliptic curve and one post quantum, or you wanna, like, not be be generic over the crypto curve, leave that out of the general architecture or repo document? Do you wanna fix the curves, or do you wanna define this abstract to the curves?

[00:45:36] Daniel Holmgren: That's a good question. I don't know that I have a super strong opinion on it. I my my sense is probably to define it abstractly, but to but I I do still think that we need to ground, like, the interoperability in in a particular set of curves. So if that looks like a registry or I don't know. I don't know exactly what that looks like.

[00:45:54] David Skenazi: David Skenazi, I didn't think I would be arguing about cryptography today. Hooray. No. You've got it exactly right. You're gonna need some level of registry for the future so you can evolve things. But every time you add a new curve, you add an opportunity for interoplimitations to no longer interoperate. Right. So you're gonna need a mandatory to implement. Like, that's pretty much what TLS did. And for this, it any any one curve you add will cause you pain and this ability to dissect. So I would say keep this number as small as possible and maybe, yeah, keep the one that's been supported thus far. Keep a post quantum one, and then have a registry for the future. Mhmm. That's your best guess at keeping it really simple and not forking the community with some implementations not being able to talk to others. Sweet. Thank you.

[00:46:44] Roman Danyliw: So

[00:46:48] Shuping Peng: Oh. Yeah. So, Eric.

[00:46:50] Eric Rescorla: Yeah. So first, I'll teach you the technical point. MLDSA is not a curve. I mean, like, you know, there there I think signature algorithm is the right level of distraction. We made this mistake in TLS of using the word curve to, like, refer a group to refer to, like, to key exchange algorithms. We thought we were doing Diffie Hellman, and, like, all the terminology is screwing now because, like, MLChem isn't isn't a group system. But, so I think that's the first point. The second point is I think David's David's totally right. We need, like, a migration story. Like, I don't quite understand, what the story is even here. Like, say we added a any new thing no matter what it is. Right? The problem is is that as as Brian is now saying, there's no negotiation. Right? And so you need to have some story about, like, how do you roll out a new algorithm in a world where you don't know what the new people support? And that's actually a much more serious problem than than, like, exactly what we choose. It's like, how do we have a story about how you make any migration at all? And I don't know if that's been discussed, but that's probably a thing we have to deal with. And and it goes two ways. Right? Because, like, you know, how do you how how do I roll out? I can't sign with a new thing until everybody supports it, and I can't, like, deprecate my old thing until until until until, like, you know, so it's, having having such a story there is, like, is is quite quite key.

[00:48:07] Daniel Holmgren: Yeah. That's a great point. And, yes, any change like this is going to take a lot of care to roll out. As I said, commits are probably the easiest thing for us to update, and these signatures are on commits. Like, right now, I think that there's there's enough coordination in the network that something like this could be rolled out with quite a bit of care, which is not to say it wouldn't be a huge pain, but I think that we could pull it off.

[00:48:29] Eric Rescorla: Yeah. I think we I mean, they they maybe now, but, like, in in if if if this really works, then it's getting much harder. Right? I mean, this is probably the most closest analogy we have here is DNS. Right? Where DNSSEC basically, like, you know, these are one way communications and, like, you know, and, like and it's been a real problem to, like, to upgrade upgrade algorithms in DNSSEC. So I think I think you're gonna need, like, duplicate signatures for a very long time you know, concurrent signatures for a very long time, and then you have rules around the manager. So, like, we need to write some migration rules at some point pretty soon.

[00:48:59] Daniel Holmgren: Yep.

[00:49:00] Shuping Peng: Okay. I'm going to lock the queue for now and let Daniel finish.

[00:49:05] Daniel Holmgren: Okay. Sweet. I had I have a few slides to get through still. So in the in the synchronization draft, there's kind of four messages that go out over the synchronization protocol. These are examples of those four messages, commit, sync, account, and identity. We have some similar changes to make to this as to the, like, the commit objects, which are basically there's some fields on these that reference DIDs, updating those to reference repos, and then just deprecating a few unused fields that were kept around for backwards compatibility purposes, two big blobs and handle on the identity event. For framing on the WebSocket, we have this, like, kind of, like, WebSocket subprotocol for events that go out over it. Currently, that looks like, two CBOR objects that both get CBOR encoded and then concatenated, which is valid CBOR. There's sort of this problem that a lot of CBOR libraries don't really know how to parse that or work with it. It's kind of led to friction for certain developers. And so one thing that we're interested in is reworking that framing protocol so that it's just a single object that goes over with the type so that these things can be processed generically and then the payload of it that can get handed off to the application. There's a problem with so, generally, these subscriptions work with monotonic sequence numbers, that a consumer can keep track of and know where they dropped off and pick back up on the stream in the same spot. The problem with this is that cursors are currently sent on the messages on the protocol. And so if you subscribe to a low throughput host, consumers may connect and then never get a cursor before the connection runs out, and then drop off and come back on and, like, kind of not know what they there's there's basically this, like, weird problem where you're not getting an update on the on the cursor if you left for a bit and then came back. And so one thing that we could do is introduce a new info type of message that basically tells a consumer what the current cursor is whenever they connect. It's helpful for low throughput hosts like this to basically get their bearings. It's very cheap to send out on on a startup of the subscription. And then there's been some discussion about maybe alternate sequence types. Like, right now, the sequence numbers are kind of arbitrary monotonic sequence numbers by the provider. That makes it very difficult to switch between two different sync providers. If one of them goes down and you wanna jump over to another one, they probably have a totally separate history of their cursor numbers. In the in past versions of the sync protocol, it's, like, very important that you never drop a message. Now this is more easily identifiable and recoverable that maybe gives us some wiggle room on this. Some services that developers run-in the network use time based cursors to allow for easy cutover rather than these kind of, like, arbitrary sequence numbers. Maybe we could introduce some sort of semantic for time based cursors into the synchronization protocol or maybe offer a way that allows both options, either the monotonic sequence numbers if you want to ensure that you never miss a message from a given host or time based if you want to be able to easily switch between two different providers. And then I'm gonna wrap up real quick because I just have one slide left, and then we can do questions. And then, another semantic that, like, we're kind of interested at. This is this has been something that's a common ask from developers is hanging on to past versions of records in an authenticated way in the repository. In these past but this would be an optional thing that you could do. So for instance, if you have, like, a, a post or something in your repository, you make an edit to it, and you wanna keep around an authenticated copy of the previous version of it. Storing it in the repository, this would require an update to the ATURI format that would basically add another component to it that is the hash of the record or the CID of the record. That would, like, structurally since the repository is lexicographically ordered, it would structurally place them adjacent to the current version of the record. This intersects with the discussion on URIs, obviously, and we can get into it there. This is probably a deeper one, but I basically wanted to signpost it and say that we're interested in supporting this. Yep. Sweet. So that's basically an overview of everything. We got a couple minutes left if people wanna hop in with any any questions or kind of just thoughts on the scope of this.

[00:53:28] Shuping Peng: Okay. Tom.

[00:53:31] Dietrich Ayala: Tom

[00:53:31] Tom: here. I did a full network study for a month to understand a product a bit better for my research group. And one of those factors also mentioned by you was the Cursor system. I wanted to look at it from different providers. But the question for me, including the blog post by Brian, was how feasible it is to still use the time based cursor for layovers between different providers as the network grows. Because, the more commits we get, the more I would assume that the time based cursor will probably differentiate between different providers.

[00:54:06] Daniel Holmgren: Yeah. That's true. I mean, that is the trade off with the time based cursor is that it it does differ between different providers. With with these fire hoses, like, generally, low latency is, like, pretty important. And so I would expect that between, like, kind of reputable providers, it's not going to differ by more than a few seconds or something like that. But, obviously, there's clock drift on machines, and, you know, they're going to be crawling these records at different rates, and there's partition different partitions in the network. So that that is the downside of it is you lose out on those guarantee the the completeness guarantee from the cursor, and you have to recover that from the data structure, basically.

[00:54:42] Roman Danyliw: Yep. Thank you.

[00:54:44] Shuping Peng: Okay. Thank you. We move on.

[00:54:57] Brian Newbold: Alright. I'm back. So the previous one, to be clear, was, that that early work is, like, pretty active. So we're really looking for for sooner like, nearer term feedback on the scoping and how to merge things. And, like, I didn't hear a ton of objections, so I'll probably go ahead with a lot of the changes I presented. This one is, like, a little bit, like, just throwing this problem out a little early. I don't in terms of timelines, I'd be kinda disappointed in myself if we didn't have something that we could bring to the group by the next ITF in San Francisco for the repo and sync stuff. But for this URI thing, I'm kind of guessing we're gonna be just getting into this or, like, trying to to brainstorm and come up you know, just starting to draft on these URI issues between now and then. So we're not I'm not trying to push, like, really hard and fast to get the consensus on this. I just wanted to bring it. It's a gnarly problem, so I wanted to bring it early and let people start thinking about it when they're eating pizza. Alright. So I'm gonna go over kind of what the current syntax is, like, what's nice about the current at URI thing, what the problem is specifically in terms of URI syntax, some possible then then one category of approach we could take here is changing. I'm kinda just replicating some of this. I wrote a blog post about this and sent a email out on the mailing list. So I'm kinda just going over a lot of this again. But some potential syntax changes we could make, maybe another option is try to get the IETF URI generic syntax rules changed. That's like moving the mountain. A lot of people are gonna not like that, but I kinda wanna poke on it a little bit and see just how big and heavy that mountain is. It seems like it's probably very big and very heavy, and we don't wanna climb up there. And then some other kinda, like, pragmatic comments on all this. Like, what is this? Regardless of what we do, does it all really even matter a little bit? And then what I think like, kind of proposing some next steps for this. So I would say if you have comments if you have meta comments on this whole thing, maybe leave those to the end of this. If you have little comments along the way, maybe also just, like, put them in the chat or something like this. I kinda wanna leave try to structure this a little more as, like, discussion at the end instead of, like, nitpicking everything little thing as we go on. So just as a reintroduction to, you know, how, the generic comment the the URI kind of flavor of URIs are usually a scheme. There's an authority section, which often has maybe some user info, a host name or an IP address, and then path segments. A little more complicated than this, but that's kind of the general flavor. The at URI like, one of the big ideas of at proto is that we put the account identifier in the authority section instead of a network location. And you resolve that account. There's another layer of re indirection where you can resolve the account identifier to a network location, the current network location. But doing things this way means the accounts can move between network locations without changing the URIs. They're like cool URLs. They don't change when you move between different locations in the network. And here's some examples of of at your eyes using the two current supported account account identifier resolution schemes, which are did PLC and did web. Some of the things that are nice about this, like, they're pretty friendly. You know, they look kinda like URLs or URIs. Like, people are pretty like, common folks are pretty familiar with these colon slash slash URLs. There's I don't think I've seen anyone with a tattoo quite yet, but I see lots of, like, stickers of these with that colon slash slash. I have one on my laptop right now. I have, like, this this magnet at home. It's kind of this thing that has some energy behind it. There's no escaping or encoding needed at all with the AT URIs right now. They are URIs, and URIs do have this generic percent encoding thing. But as we've structured and kind of specified things right now, you never need to do that. You never need to do a pass to remove percent encoding because there's nothing that needs to be percent encoded in theory in the current URIs. So they're nice and normalized. There's no optional slashes at the end. We're kind of strict about how all that works. So you just have these nice, clean norm like, they're kind of always normalized URIs. And the account identifiers are just pasted right in there verbatim. It's very easy to see that that's what it is. You don't have two different forms of it. Developers don't get confused. Users don't get confused. You can just parse them in and out. And the current syntax has all these properties. It has all this great stuff, and it's deployed. There's just billions of these strings out there in these content address records. So there's just, like, a bunch of these strings that look like that out there that we're gonna need to deal with going forward. The problem is that you can't this just doesn't work with the RC thirty nine eighty six URI syntax rules. When I went you know, we went back and reread it really carefully. If you use the slash slash, you have to follow the authority rules that it has to match, like, a registerable name or an I p v four or an I p v six in there. And that's, like, kind of doubled down on in the AYANA URI registry rules, which specifically mentioned this. It's like, if you didn't do the slash slash, you don't have to follow those rules. But if you did do the slash slash, you gotta do that part. It's like, oh, damn. Sorry. This was kind of my fault that we ended up in this position. I reread these documents over and over again and somehow convinced myself we didn't need to do it, but now it's very clear to me when I go back. It's not like there's an error in the documents. You just need to do it this way. And this is already starting the as people are trying to use these URIs, there's no problems within the App Proto ecosystem right now. Like, people mostly just write their own little URI parser for these in their kind of, like, SDKs for App Proto. Everything works well within the repositories. They work fine in the bodies of API requests and stuff like that. But if you start trying to use these as your rise in other contexts, it's starting to cause problems. A good example is people trying to put these in like, link rel alt HTML metadata, and then suddenly your HTML doesn't strictly validate because there's expected to be a URI in there, and it's not a URI. And we want people to take these cool URIs that don't change and spread them all around the web, spread them all around in other places. And if they're not strictly, you know, ITF URIs, that's causing problems. So a couple ways to you know, a couple ideas we've had, some some changes we could make to the URI syntax to get around this, to get them back onto the valid path. One is to change don't make it colon slash slash. We could do three. You know, three is better. Have three slashes. You have an empty authority scheme. It's kind of subverting the URI concept, but it kinda works. You lose this idiomatic prefix. You lose some, like, style points or, like like, philosophy. This feels a little like you're we're doing something a little bit dirty, but you don't have any escaping or coding. And the the identifiers are just right in there. It's just a single character change. Another approach would be to remove the slash slash. You have a t colon, but now you're still like, oh, what is this? Is this a URI? Is it a did? It's, like, a little weird. I can tell everyone is chatting seriously. There's like an energy in the room. Another approach would be to do the percent encoding. So just take those comments, encode them. No one likes this. I'm glad other people you I point that you get both you have to do with escaping. You maybe need to double escape. If you put these in a URI query parameter, you're gonna escape the percents then. Right? And you don't have this visibly it's not as nice when you look at it. You don't it's not obviously your DID when you see it there. So you lose a bunch of stuff when you do that. We could come up with some clever transforms. Maybe we map these down. Like, this is probably, again, kinda doing it dirty. People won't like this because what's there's what's dot did? PLC dot did. Where did that come from? That's not a TLD. Are we reserving these? Are we trying to get a special, TLD like dot onion? Put you know, this is just gonna be more processed. It's not great. Maybe we do them with dashes or underscores or other things that make them clearly not valid local host names. Like, it's not a host name, but it could be a valid DNS name. It could fit the syntax. Anyways, this is a approach we could take. The big gnarly one would be trying to get the the URI rules changed. So we could try to restructure things where the authority section still has some character set limits. That would be basically the same character set limits, but there wouldn't be the substructure in the same way. So in an in a in a SDK or a library, you'd parse out. You'd get a string for the authority and optional other stuff or something special for DIDs, but we're trying to not bake in DIDs. You know, we wanna leave these, resolvable account identifiers a little bit more flexible. This is just gonna be hard. I kind of wanna at least poke a little bit more, like, get a little bit more feedback, including this presentation in this room about this, but I don't think we can bet the farm on it kind of. And even if we got the, you know, the RFCs changed, like, there's just a ton of code out there in the world. It's a hard it's you know, getting RFCs changed doesn't change all that code that's out there. There's still gonna be problems. One approach is just don't call them URIs. I think people have kind of indicated of this. Like, they look like URIs. They walk like URIs, but we could just call them identifiers of another type. This feels a little off to me. Everyone's gonna be confused. People are gonna try to use them as URIs, etcetera, etcetera. And then I wanna just kind of broach. I'm per at least personally, like, very confident. We need to at least specify the current string syntax, acknowledge that they're not valid URIs, but they're out there. Proto implementations are gonna need to parse these out of records that are content addressed and around there. So there's gonna need to be this kind of, like, legacy support for the existing strings. And if you need to do that legacy support, will folks just keep doing this? And so that that, mean, that almost kind of points at this don't call them URIs thing, which is, are there always gonna be two versions of this, and how do we stop people from putting them out? Do we try to come up with a transition where we really go and say, oh, no. Stop omitting those. You're really not supposed to do it that way. You're supposed to only use the, like, the correct UI syntax. So to me so okay. So this is the thing. We've got time. There's no, you know, totally hanging sword over our hands, but we do have do wanna get some deadlines for wrapping up this charter. We're gonna have to get to it at some point. My kind of proposed next steps here are take the existing syntax rules, which are invalid, to be clear, but just write that up in an IDE so we have something very informal that we can point people to and talk about the problem. Probably send it over to the URI. Send this whole thing over to the URI review list as a, hey. Look at this big mess. What are your what are your thoughts? I just think that's, like, a thoughtful group of people. I don't wanna waste their time or say, hey. This is your problem. Like, tell us what to do. But I wanna if we're gonna do it, like, I think it should be, like, at least a clean maybe not my personal blog post or something that we send over. Like, send an idea over. Say, like, here's the dilemma we have. Here's what we're thinking. And my guess is we're gonna have this conversation again, but maybe we can but maybe by the next one, we can be getting more like, okay. Let's be kinda driving towards consensus. So sorry again to everyone. Thank you for everyone's input on this, and we can go to the queue.

[01:06:44] Shuping Peng: Ted.

[01:06:44] Ted Hardie: Hey, Brian. My name is Ted. How are you?

[01:06:49] Eric Rescorla: Great.

[01:06:50] Ted Hardie: I was gonna say no relevant affiliation, but actually author one of the authors of the registration document. One of the principal things that registration document does is trying to make sure that there are no collisions. And I really appreciate your coming early and getting the feedback from the list there to make sure that you are picking a string so that there's no collision between what you're using and something else. That's something you did well, and I wanted to call it out early because it's not really the best thing that happened after that. And your representation of the generic syntax of URIs in this presentation is actually wrong. Because what you're focused on is URIs which use an authority, and I think you've got some some things you can do here that are relatively simple if you decide, hey, it turns out I don't love slashes so much. You can take very much what you've got now and do a t colon, the did, pl c colon, and go from there, and it will look exactly the same as what you have here, and it will be a perfectly valid URI because you will instead of having an authority there, you will be having, your BNF will say that the set of things which can appear here are valid DID methods, or valid DID methods, plus registrations which occur in the following IANA registry. Yeah, yeah. So that's a valid URI, and it's a very simple transform, what you have now. You're gonna miss the cute slashes, and I think you should just put them away and say, they were great. They were for they're the slashes of my youth. I'll always miss them. But it's time to become a man. And or a woman if you are one, or a non binary pal if you're a non binary pal. Any of those, but we're we're going to just move into the actual URI syntax here. All of the other things that you've described other than putting away the slashes are hacks. And they they may be valid hacks if you really want to go there, and yet, can they leave the slashes behind? We can help you find a hack. But that's the simple thing to do. I I will also point out here that, one, congratulations, you're in the IETF. You have a charter. Changing the URI syntax ain't in your charter, so you can they do that in this group. If you decide you want to change the the URI syntax, you get to go and ask dispatch where to go. I don't suggest you do that without having had calming beverages beforehand and a large supply of them after because I can tell you that they're not gonna tell you to go any place fun. I I really do think that we can get you someplace sensible here if you're willing to just say, I'm gonna bite the bullet and go straight to URI syntax. If you're not willing to do that and you wanna call these something other than URI, life goes on. They're the the colon and slash are not particularly special in the world. But you're gonna have to figure out some out of band way to signpost that very strictly, because otherwise, you're gonna run into parsers that think, hey. This looks like a URI, and I'm gonna behave that way. We we actually just recently processed a document that's finding ways for researchers to reference URIs that are to malware sites or similar so that they are not automatically parsed by parsers. And it is complicated to just get all those parsers out there that want to see anything with a colon and a couple of slashes as a URI from doing the their their their erstwhile best. And that signposting is gonna be almost as hard as just going straight to the URI, the syntax that actually works. So happy to help with all of this. Have a great day.

[01:10:39] Brian Newbold: Thanks for your patience. Hi.

[01:10:43] David Skenazi: I didn't get the memo that Thursday open mic night at the comedy club got moved, but I am excited. Hi, everyone. My name's David Skenazi. I'm a veteran of the r c sixty eight seventy four wars. For most of you who is that don't remember, that was the last time we tried and failed to change the URI syntax. The idea was, oh, I p v six addresses are great, but we wanna put the scope ID for link local addresses in there. I fail to find words that are polite to describe how that went down. It turned into a fight between the a p v six side of the house and the HTTP side of the house. And then we realized that actually URIs don't matter because the web working group ran away with URLs, and that's what the web uses. And, anyway, I I could go on and on. Ted did a great job explaining the world of hell that you would find yourself in if you tried to do that. Ask me over drinks about the small portion of it I experienced. I think Ted summarized it well. If you don't wanna change what you have, call it a string. If you're willing to change, change it somehow. I think, yeah, dropping the slashes is the simplest change. But, you know, if you're changing, might as well change it to the ideal thing in what you think it is. But, yeah, the dog, please, for your own sake.

[01:12:08] Dietrich Ayala: But you did

[01:12:09] Brian Newbold: it, so there's a chance.

[01:12:10] David Skenazi: Oh, no. It wasn't. I was a bi I oh, no. We didn't. It got rolled back. I was a bystander. I just have scars, and we accomplished nothing.

[01:12:21] Ben Goering: Hi.

[01:12:27] Justin Richer: Justin Richer, welcome to the ITF, guys. Just I do wanna call back to so I I helped chair the boff for this, and I do wanna call back to when you come to the IETF, they will call your baby ugly. I'm really glad that you realized this particular dimension of baby ugliness early because, yeah, now is the now is the time to dig in and figure out what direction that you wanna take. I don't think there is a good answer here. The worst answer is hoping that you can change URI. Because and I was I was gonna bring this up if David had an what WG effectively owns that as much as the IETF doesn't want to admit it. Sorry, guys. And so although those are living documents, maybe you could get them to change it. Who knows? No. Don't do that. That was a joke for the record. Regardless, yeah, I think that it's it's ultimately going to be a syntax or a semantics change here regardless. I don't think it can just it obviously can't live how it is because it it doesn't line up with the definitions. And so that since you're already looking at places where backwards compatibility and mapping to backwards compatibility is gonna have to be an explicit thing, I don't think that's the worst thing in the world. You know? I think that you can either decide that, hey. It's a syntax change and we have a transform for syntax change and we realize what that means, or it's a semantics change and you pretend it's not a URI and then you go out of your way to warn everybody that no, no, it's not really a real URI, so don't treat it like one. It just starts with something that smells like one, and you do run into all the problems that that was talking about earlier. So but that is a choice. So I don't think there's an easy way out, but here we are.

[01:14:24] Shuping Peng: Okay. Thank you. So please keep short. We Sorry.

[01:14:28] Justin Richer: I'll try

[01:14:28] Ben Goering: to keep jokes out of this. So, yeah, the charters talks about how cryptographic verification ensures existential and forgeability of repository data. And I think, like, that's that's that's what we want. That's what we should focus on. And I'm hearing different things. I'm like, it's hard to migrate an existing network, we need to document how it currently works, maybe even verify whether the current network provides that. But this was a whole discussion of, like there's several discussions here of, like, things we could do, and it I just think a lot of that should come later once we have work items in the group and we're even so that we can focus on, like, the high priority here of of the cryptographic guarantees that are in the charter. And then I guess the question is, what things by the next IETF will be called for adoption? Like, what work are you gonna do before the next IETF so that this all this feedback is going into working group items and not some company's thing. That's a question. Are there any documents that you'll you'll do before

[01:15:27] Brian Newbold: the next? My feeling like, I'm pushing on the repo and sync docs, then I kinda expect those to go forward first. In this URI thing, I'm just kinda planting the conversation going back.

[01:15:38] Marcus Sabadello: Hi. I'm Marcus Avotelo. I'm one of the coauthors of the w three c DID standard. As a as a DID person, I love the idea of building DIDs fundamentally into URI syntax. But as others have pointed out, I think it's might be difficult. Could you go back to the slide maybe where you have the three triple slash option and the other option? Wanted to point out two things here. First of all, the triple slash option, it's not that weird. Think there are other URI schemes that do that. For example, file URI scheme, you can have file colon triple slash, which is perfectly valid and well understood. And, the option without any slash, I think there is one issue here that hasn't been discussed yet in the mailing list, which is that in the DID standard, we also have a concept of a DID URL where the DID can have a path. And if you I think if you use the syntax without any slash here and then you have a path, then semantically, that path might be considered to be owned by the DID and by the DID method, which could mean that the path might be dereferenced in the DID method specific way, which is I think not what we want here. Whereas if you have the triple slash, then the DID is one segment in the URI and the path would be fully controlled by the outer URI scheme, ATE, which I think is what we want here. So that's what I want to I wanted to point out. There is also the option with a single slash here, right, that hasn't been discussed. That might Yeah.

[01:17:04] Brian Newbold: I can't remember the the single slash might invoke the the parsing. I do some other people at other points kind of mentioned, like, just don't even just use the did path thing. I'm kind of against that. I think it should be clear that it's we're talking about AT, and I'm pretty sure we could come up with a URI syntax saying that it needs to be the account identifier without any other path. So I think we could resolve that part.

[01:17:27] Phillip Hallam-Baker: So back when I Philharm Baker. Back when I was young, thirty five years ago, when we added the double slash to the HTTP and FTP methods, the reason that they're there is so that you can do relative directions so that when you've got one URI and you give a partial one later on in a document, that it can say, here's a new authority. Here's a new here's a thing that starts in the root of root path in that authority. And so that's why they were added in. So I really don't like the triple slash any more than the double. I think that that gets you into saying, this is a URI that you can interpret as a relative URI to something else. You can't. So let's just get rid of the slashes altogether. And I think that the the last one, a t colon, actually is gonna be the one that you want because what you have some there is something that means the same thing irregardless of where you are interpreting it from. So that's the correct approach, I think.

[01:18:48] Shuping Peng: Okay. Sorry, Jan. Very short one.

[01:18:51] David Skenazi: Of course. Janem dot w social. I just wanted to point out, you mentioned earlier that the parsing doesn't need to be care about percents, character encoding and stuff like that, and that's very nice. But if you call this not a URI, I think I'm actually at the liberty of percent encoding every single character in my data, and you should be able to parse that anyway. Just wanted to point it out.

[01:19:11] Shuping Peng: Thank you. K.

[01:19:15] Brian Newbold: Alright. Thanks.

[01:19:31] Shuping Peng: Eli. Yeah.

[01:19:35] Eli Mallon: Hello. Can everyone still hear me?

[01:19:37] Shuping Peng: I have the control.

[01:19:38] Eli Mallon: Okay. Thank you. Okay. Excellent. Excellent. Thank you very much for the invitation. Today, I'm going to be talking about at proto over media over quick transport, which is some work that we just concluded. So my name is Eli Mallon. I am the founder and CEO of Streamplace. If Blue Sky is using the app protocol to make a Twitter analog, we are using it to make Twitch and YouTube do a first approximation. I was just minding my own business last March. I would been running a streaming platform that used primarily RTMP ingest and WebRTC playback. And I've been pretty happy with how that's been going overall. And all of my video engineer friends were getting more and more excited about this thing called Media Overquick. And I looked into it a little bit, but I was like, oh, I I don't know. I mean, we still need to support these other protocols. It's not supported everywhere yet. So let's let's get good at these other things first. And then in March, two Cisco engineers dropped onto the IETF mailing list this draft for doing at proto data, specifically the relay role in the at proto ecosystem using at media over quick. And nothing derailed my technical roadmap in a way that nothing really really has in recent memory because it's like, oh, I can just start doing this one protocol and use that for all of our media things and use that for all of our app proto data. You can even synchronize it to the browser with web transport. So had to throw away the whole technical road map and and we're doing this instead. So all we're talking about here is currently in sort of the general generalized AppProto architecture, you have a connection between a PDS and a relay that goes over WebSocket. Relays aggregate all of the different PDSs and go to app servers, formerly app views. And then apps app servers and the and the clients can, course, use whatever protocols they want to sync. But a common choice is to use what's called the jet stream, is like a simplified version of the app proto c or relay fire hose. And theoretically, we could use media over quick transport for all of these different roles. The work that we've just completed is and to have in production in Streamplace is between the relay. So we've implemented a relay server and then Streamplays as a as a consumer of that, as a as a client in there. I'm gonna do there's lots of good overviews on what Media Over Quick is, so I'm just gonna try and give a very, quick sort of what is mock for at proto engineers here. It stands for media over quick transport. Quick is, of course, the UDP protocol that was designed to be the basis of HTTP three. It's being standardized down the hall. The authors of this draft would be in the room actually, except we happen to have it scheduled. So for for future meetings, I've asked if these two meetings can be listed as a conflict with each other. The the fundamental basis here is it's a big global caching, a public published subscribe system, and it was created to replace big h t t p c d n's. So video technology still has it it still looks a lot like putting a caching NGINX or or Cloudflare or something in front of a media server and caching all of the media that that moves through it. Right? And there's been a lot of like Apple's LLHLS proposal and a c math. There's been there's been a few attempts to modernize this a little bit, but this is like behind the scenes. Most live video you watch is still, like, making HTTP requests and and caching at that layer. So everybody's sick of that, which is why we've been able to get such consensus around media over QUIC. And, yeah, like I said, all of the video engineers I know are very, excited about this. And this has a lot of promise to replace a great deal of inadequate protocols that we currently use for media broadcast. So why is it a good fit? For at proto data, the the biggest one is everything right now is handled over WebSockets. WebSockets map to a single TCP connection. I from just what I was measuring yesterday, I was getting about 17 megabits per second over the fire hose. We are all, of course, in this room hoping at Proto, we'll scale 100 x, hopefully 1,000 x someday, and you will just absolutely run into scalability issues attempting to send all of that data over a single WebSocket. And so, you know, if you're you need to go find somebody that knows how to broadcast lots of data at the gigabit and terabit scales, you're going to talk to video engineers. So that, you know, on in one sense, this is just a replacement, sort of a drop in replacement for WebSockets with much higher throughput, lower latency, and better caching parameters. But to me, the real headline here, as I mentioned on the mailing list, is that we can doing this would allow us to use commodity media over quick relays. And we've tested this successfully in the in the Atmoq project. So currently, the relay architecture in AdProto looks something like this. You've got sort of the root relay. And then if you're going to fan it out, Blue Sky has this piece of software called Rainbow that sits in front of them and and sort of handles fanning out that to to lots and lots of different servers. Rather than doing that with bespoke architecture in the media over quick world, we could do this with commodity mocked relays that are being worked worked on at Akamai, Cisco, Cloudflare, a lot of giant companies. So it really, like, just sort of slots in at Proto here and and gives us this really, really good gigabit terabit scale scaling without having to design all of our own software and realize and that sort of thing. So this is what I think is the the really exciting thing for me. So what do we have right now? I'll and I'll get into sort of the protocol considerations in a moment. But so we didn't implement a full relay. It's sort of a proof of concept sitting in front of an actual relay that aggregates stuff. Yes. Thank you for clarifying, Brian. There's not a single source of truth relay in in in AtProto. I just mean to say this is if you were operating one of many AtProto relays and you wanted to scale it up, you would currently use this sort of bespoke software that that goes in front of it. There's not like a single signer or something like that. We have a seventy two hour replay window via group IDs. I'll get into the sort of primitives of this in a moment. There's a Rust server, a Rust go and TypeScript libraries. And we're live in production right now with the the stream place dot network relay that we're that we're operating there, which you can try right now. I spun it up in Frankfurt for

[01:27:02] Daniel Holmgren: this just to see if

[01:27:03] Eli Mallon: I could really blow everybody in in Vienna that way with the the low latency. So sort of mapping some things onto here and and some of the particular capabilities we get from this. So events, the the standard events in the app proto fire hose map onto objects. These would be frames in the in the media ecosystem. One thing we don't have in existing realize that I've sort of added on here. So users in this become tracks. So just like you can request a audio in English or audio in French or something like that, You can now request different users in the same way. So this allows you to listen to if I only care about 50 users events at any point. So I can I can request that directly from the Relay and get it very efficiently? We have this existing concept of a cursor. Brian, I think, actually sort of summarized some of the the state the sort of ambiguous state of the cursor pretty well. I won't get into that. But one thing we don't have is groups. So it would be a little bit for for maximum efficiency on this, you probably don't want to individually address every single event from every single user. So you instead bundle them up into groups in the media world. This is like one group of pictures, key frame intervals, like a key frame followed by a bunch of iframes. And, yeah, that's like the the caching unit basis in this. So this is one way the the sequence numbering is one thing that sort of changed in in this here. And then, yeah, the subscribe repos synchronization mechanism becomes the subscribe live edge at mock request or media for quick request, I should say. So this is the IPF. The we're here to standardize. So the coming out of this work, there were a couple different pieces that I think could be very helpful to standardize and get a handle on. The first one is whether this, and and and one reason I do think the sort of defining the relay role in the ITF work is that it would allow us to so when we're talking about broadcasting over commodity media, over quick relays, we're talking about unidirectional broadcast without sort of client negotiations. So the question is what format do we send down that wire. Right? There until recently, there was just one sort of Firehose format, which was this sort of two concatenated CBOR objects with a header and a payload. Recently, Blue Sky introduced those like a a that that became CBOR v zero. There's also CBOR v one, which puts them both in one CBOR object as well as a canonical JSON format. I think So that sort of precludes unless I rebroadcast, you know, I can broadcast everything three times in three different formats. But because we're talking about unidirectional broadcast, I I I have to sort of choose which one we're doing there. So, you know, the the analogy is like, if we were going to send all of this data via satellite or something like that, what format would we wanna use? And I'm happy to use a variety of formats, but what we can't do is is two way negotiation as to what format we we want to use there. Right? This is strictly unidirectional broadcast. And so if we think if we want this to be like a drop in replacement for WebSockets and and we think this is this is valuable, We're all of these apps are going to have to speak whatever format this relay speaks anyway. So I don't know. Maybe maybe we drop this negotiation piece. The other one here is this sort of comes down to so I I did went out and did a big survey. There are, to my knowledge, five complete implementations of an app proto relay right now. And all of them have slightly different cases on data handling. So there's sort of this sort of notion that relays should be generally permissive of invalid data, data that they don't immediately recognize because we want them to be future compatible. We might add more stuff, more fields to the to at Proto in the future. And we don't want relays to be, you know, immediately breaking because that we we not everything is in there. But so what I've I've produced here, you can take a look at the link, is a series of tests that run against all four all five Relay implementations.

[01:31:37] Ori Steele: And all of

[01:31:38] Eli Mallon: them sort of disagree on what should and shouldn't get passed through from PDSs to everybody else. And when things do fail, the behavior on that failure is different. Indigo tends to dump the entire incoming connection from the PDS. There are, yeah, a variety of different ways we can go in there. My like loosely held advocacy here would be that we should be strict around passing DAG CBOR, also called Drizzle, the the IPFS, Dazzle Projects, the Drizzle specification, which is, like, specifying the CBOR encoding, the the key ordering, all of these sorts of things. Like, let's let's all speak that because then we're speaking the base format, and crucially, we can content address it well. But and then but but be tolerant to sort of unknown app protosemantics. Right? If somebody includes other like dollar prefixed fields or that sort of thing, pass those through but don't let just like arbitrary c port through. There's a lot of ways to do it, but it's a it's a very inconsistent patchwork in the real world. So that seems like something we could be standardizing. I've just got one more slide I'll go through and then call you up, Amelia. I just wanna say thanks to Suhas and Fluffy Jangs for for dropping the draft and completely derailing my technical road map, to Richard Barnes for getting us all connected, and to Brian Newbold for soliciting this presentation.

[01:33:17] Shuping Peng: Okay. Hi, Melia.

[01:33:19] Amelia: Yeah. So on the previous slide that you had there, Eli, I'm wondering, can this conformance test suite be used as a input to create issues for each of these weird cases and where things are not handled consistently and used as an input document on, I think, the synchronization specification or Internet draft? Sorry. Could we use it as an input document for that? And create issues and issue tracker for the synchronization spec.

[01:33:53] Eli Mallon: I think that's a great idea. I mean, yeah, we probably have to start with, like, identify every disagreement and and decide what we want the the sort of approved behavior to be. But yeah. Yeah. I think that's definitely what the process should look like. Something we haven't finished quite yet, but I I want to have off of this would be a, like, a fully currently, this is instrument this this was produced by sort of instrumenting all of these different PDS all of these different relays. I I would like to have, like, just a fully external, like, throw any relay at a test suite and get a result.

[01:34:27] Amelia: That would that would make sense. And, yeah, I I I know for, like, w three c k w three c standards groups, we have these sort of conformance test suites. I'm not sure in w three c context what that is. So

[01:34:46] Eli Mallon: Right on. Cool. Thanks, everybody.

[01:34:50] Shuping Peng: K. Thank you. Okay. Now you have the control. Oops. Sorry.

[01:35:07] Phil Feigl: I think the order changed. So these are my slides.

[01:35:13] Shuping Peng: You wanted the control.

[01:35:16] Phil Feigl: Sorry. Sorry. These aren't my slides. I think the order changed. It's Floats (Juan Caballero)

[01:35:21] Shuping Peng: Alright. Okay.

[01:35:46] Phil Feigl: Well, I'll start, while the slides load. I'm Fig or Phil from, Microcosm. This is my second ATF first time presenting. Microcosm runs protocol infrastructure and data services used by hundreds of AppProto apps, and we have implementations of both spec drafts running in production. I wanna bring up a problem in the current repository spec draft that I think is gonna force a decision we need to make, which is we actually need a way to signal whether our repository archives blocks our stream ordered or not. If the order if the slides are stuck on this one and

[01:36:26] Roman Danyliw: We're working on it. Just continue.

[01:36:29] Phil Feigl: Mhmm. Okay. So we have this spec should in the repository spec that says, yeah, there we go. So producers, so that's PDSs, should emit blocks in preorder traversal, and parsers must tolerate other orderings. The block ordering section then makes the claim that this enables parsers to avoid buffering most of the blocks when they're stream ordered. I'm certain that this claim under the parsers tolerant must without an explicit signal is false. So this actually matters quite a lot because buffering blocks for nonstream ordered repositories is a huge driver of current backfill design, which means that it affects almost every app that wants to build on the AT specs at some level. In particular, it affects Microcosm where we target full network services that scale down to run on small and affordable service. So this is a 1% sample of repository archives on the live network bucketed by repository size. So blue is the number of repositories in that bucket, and red is the sum of the sizes of those, repositories in that bucket. So the total data weight on the network represented by each bucket of repository size. So even though almost all repositories on the network are small, the blue bars, first two buckets under one megabyte is most network. The bulk of the actual total data in the network resides in risk repositories, which are large. So 80% of all data in the network is in the repositories that are over one megabyte, And that's where stream processing starts to matter. So this really determines how you do things like backfill even though it's a small minority of total repos. So the repository holds key value data in a Merkle search tree, MST. The tree's nodes hold keys with links to records and links to child nodes. All links are content addressed hashes of their target. So records are encoded to hashable bytes with deterministic DAZZLE seaboard, and the MST nodes themselves have their own DAZZLE seaboard encoding, which again is what you hash over to get the link to that content. So everything is bytes identified and linked by hash over those bytes. And their archive format, which is a CAR file, is a flat series of pairs. So the hash of some object bytes followed by those bytes, hash of the next object's bytes followed by those house bytes plus some per block framing. If you read these pairs into a big map in memory where you key those hash links to the object bytes as the values, then walking over the tree to traverse the repo is just a hash map lookup to follow each link. This is what you have to do if the blocks in the arc archive are in some random unordered order. So preorder traversal block ordering is just putting the serialized blocks in the archive into order for depth first walk over the tree. So depth first walk over the tree visits entries in key order, and this is usually how repositories on the network are processed. So if you're stream processing this archive, which in this case is preordered at the at the bottom example, you find each block just in time as you read through the archive. The next link block that you want is the next block that appears in the archive. So you don't need the in memory map. You get to release from memory each record block after you process it, and every MST node after it and all its children have been traversed. So this is great. The problem is that records are linked by the hash of their contents. So if two records have identical contents, then they have the same hash link from their different keys. They are the same record as far as content addressing is concerned. So this is fine and legal and good. There's no uniqueness constraint on record values within a repository. It would be weird if there was one. With stream ordering, that means you have to serialize that same record multiple times into the archive. So once for each time it appears under a key. So what if you're a stream processing parser? You see record one under key one and you process and discard it, and then you reach where key two's record block should be and it has the same link. What if that record content doesn't appear again in the archive for that second key? An archive that doesn't include the record a second time is still valid. It's just not stream ordered. So how can you be a streaming parser that's tolerant of this block ordering? You have to keep record one around the first time you see it so that if it's referenced again, you can look it up. And in fact, you don't know in advance if you're gonna need any record multiple times, so you actually have to keep every previous record around in memory keyed by its hash link, the whole memory buffering thing that we're trying to avoid with hash ordering. So the whole point of stream ordering is defeated under the current spec. This isn't theoretical. So today, 0.4%. Again, this is about a 1% sample of the network. Of all repositories have records that are referenced from multiple keys. Most of these are not sort of, like, essential to the data they're trying to represent. They're more like extensive how records are created, but it can be invalid to to do that anyway, and that's kind of a problem in the spec regardless. So the fix for all this is really simple. We just need a way for a repository archive to signal to a parser that it's stream ordered or not. It cannot be implicit. The missing block in the earlier example would be would have been an invalid serialization of the repository if the archive was claiming to be stream ordered, and then you'd be able to reject it as a parser. And if the archive hadn't claimed to be stream ordered, then you would be buffering all the blocks anyway and you wouldn't have a problem. So I hope we can agree that this is a problem or maybe clarify to get some alignment on that. And these are and then I and then the challenge is how do we signal this, like, stream ordered or not problem. And these are my initial ideas. Basically, number one is we could just say all repositories must be serialized in stream ordered or, like, preorder block traversal, just not accept arbitrary block orderings, then we would have no problem. We could the next three are sort of, like, ways that early on in the archive, we could put some sort of flag. So there's a roots array in the header that we could use the second slot for that's currently unused as a sort of signal. We could put something in the commit object. It's a bit weird because it's not really part of the commit, but it would work. And the commit object has to come first in a stream order repo. We could use the first block of the archive as a sort of magic signal that this is stream ordered or not or something gross and implicit, like the commit object appearing first means it is stream ordered. We could do it over transport, like a media type parameter, which then is lost when it's at rest. But for the purpose of, like, backfill in most use cases, I think it would effectively solve this problem. So that's all I have for slides. Brian, you've got your hand up.

[01:44:29] Brian Newbold: Brian from Blue Sky. I might have missed one or two things. So I'm sorry. I had to run out really briefly. I do agree with the duplicate record, and I could be misremembering, but I kinda thought we'd that could be the crux of all this, but I thought that we'd required that, that if there's duplicate records, you have to include the So it's I guess it doesn't matter.

[01:44:50] Phil Feigl: Yeah. It's required if you're producing a stream ordered car

[01:44:54] Brian Newbold: Yeah.

[01:44:54] Phil Feigl: But it's not required if you're not producing a stream ordered car, I don't think, unless I misread that part. If it is, I think that would be onerous, but I guess maybe that would work.

[01:45:05] Brian Newbold: We have I think we like, my colleague has implemented in Go an implementation. The way so the way we did it is we optimistically assume the car is an ordered format. And so you get it, you can start reading it off the wire and go. And if you don't and but you always have to expect the next block. And so when you read off the car, your like, the API's, like, peak kind of. And if you got what you expected, you parse it, and then you drop it when you're done. And if you don't get it, the reader caches in memory until you get the block you expected. And that what we found so that's basically, like, doing this optimistically. And we found that, like, dramatically reduces memory consumption. It doesn't guarantee that your memory consumption is low. So

[01:45:51] Phil Feigl: the problem with that is that these 0.4% of repositories could be serialized in a way that your parser would break on them unless you require that the block is duplicated even if it's unordered. If you're optimistically discarding records, because so far everything has been appearing in order, then when you encounter one later that references a record you previously discarded, then your parser blows up.

[01:46:16] Brian Newbold: Okay. I'll have to dig dig back into it. Sorry. Okay.

[01:46:20] Phil Feigl: Yeah, like, I I know Mary implemented, stream parsing in at cutes parser, but buffers all blocks for this reason as well. And I, yeah, in my implementation of repo stream, I just dropped stream parsing for now until this is resolved.

[01:46:41] Shuping Peng: Okay. Thank you.

[01:47:07] Dietrich Ayala: Hello.

[01:47:09] Shuping Peng: Yeah. You have this.

[01:47:11] Dietrich Ayala: Oh, yeah. Okay. That's fair. So I'm presenting some slides. My colleague at the IPFS Foundation, Volker Misha, made. It's all his research. The links from the slides go to the repo and the blog post. So if you have if you wanna double check the work or or dig into the details, I recommend just downloading the PDF and clicking the links. So there's there's been some references already in the previous presentations to the load bearing JSON seabore mapping. It's an appendix now. I I'm not really too invested in which of these things become Internet drafts of the working group or which become work items, but I do think the tile map that had seaboard and JSON grayed out is a little optimistic. There's there's it's a profile of JSON and it's a profile of CBOR. So the profile of CBOR, we that the a t proto inherited from the IPFS foundation prior art has a way to do floats. I triple e only floats. And the argument here is that the the whole system, the JSON profile should also have floats. The currently JSON arbitrary JSON can't be embedded in AT records because the JSON profile doesn't allow it and a huge percentage of the JSON out there in the world that already exists, that's already being pumped around in various open data, scientific data, GeoJSON has floats in it. JSON with floats. So yeah, the presentation here, I'll try to go through it quickly to leave more time for q and a. So the way people are using floats in net new data in in data AT Proton native data, people are encoding floats as strings when they need floats. You can do that. You could your application logic can mark something as as a string that's actually a float and you can turn it into a float when you get it because you are expecting that because of the sort of lexicon, the schema, the JSON schema stuff. So Eli, who presented a little bit before in in his survey of what relays are actually doing in the wild, found a non zero percentage of JSON contain or a seabore containing floats being relayed around in the wild. We're not sure exactly how it got there. Some people are forking the PDS and or forking the relay disabling validation. Some people are putting floats in, you know, a roundabout way because they maybe it doesn't net new data perhaps. I don't know. So for whatever reason, people have been doing this sort of in user space. And the argument here is that it's safer and maybe in some cases cheaper to just let let floats exist at the JSON level since they're already allowed by the CBOR profile. There's they're already used in in IPFS prior art. And when digging into the details of this, Volker found that many many most JSON parsers and native data formats in all the major languages, the the 19 languages he surveyed, are already translating JSON numbers that look like floats into floats of the native data types of the language in which you're parsing JSON into. So just inheriting that convention which seems near universal could get us there an easier way and it could just be native everywhere instead of having to do it in in user land. Yeah. That's what I meant by user space. So for people that haven't dug into our cookie little CBOR profile, Integers are minimally encoded, smallest CBOR type that it'll fit into. We use floats the opposite way. We maximally encode floats, always 64 byte. 64 bit I triple e doubles. If if you really wanna get into the details,

[01:52:00] Brian Newbold: negative 0.0

[01:52:01] Dietrich Ayala: needs to be removed. The the requirement here, think, if if you just zoom out a little, any JSON has to translate to exactly one CBOR so that that's the CBOR you sign so that JSON will always verify after converting to CBOR. And yeah, there there has to be one CBOR representation for any possible JSON. JSON

[01:52:24] Brian Newbold: has a

[01:52:25] Dietrich Ayala: little ambiguity. There's multiple semantically equivalent ways to express a float or an int as a JSON number. So we're arguing to just inherit this convention that parsers do in every major language, which is that if a JSON number has a decimal point in it or an e, like one e five, that's probably a float. And, you know, just convert to at proto specific seaboard that way instead of always doing it off protocol, user space, or at the wrong layer. Yeah. It is kind of this simple. It is kind of like if you see a period or an e, it is a float. Even though JSON number has one one type, most parsers sniff for the, like, subtype for the implicit type. And, yeah, the only tricky bit is if it's a float in CBOR, when you kick it back to JSON, just put a point zero at the end. Lots of parsers already do this. This is sort of common sensical. That's what CBOR working group, I hope, would recommend. So, yeah, if you click these links, you can see the nitty gritty details and test vectors of all the major parsers. Yeah. And yeah. There's a bit of a there's there's always some wiggly bits on how differently different what JSON you get back coming back from CBOR, but that's sort of an inherent limitation. And you can always round trip and show that to the user to canonicalize. Like in in user space, you can round trip to get one JSON representation of the multiple JSON ways of representing a given integer or float. But yeah. And as mentioned, the, you know, the the content identifiers that construct a CAR file already have to add a custom parsing step in any language. So in the extreme case, in the, like, one or two languages Fulcra found where major parsers don't automatically do this. You could sniff at that same stage where you're looking for byte strings and figuring out if they're CIDs or not. Yeah. I think I should leave seven minutes. And thanks to the chairs and the group.

[01:55:09] Shuping Peng: Thank you. Discuss. No. Brian? You got

[01:55:16] Brian Newbold: Brian Newbold from Blue Sky. So thanks both for you for presenting and for Volker and other folks folks that have, like, raised this on the mailing list and done all this research. Volker, in particular, has done, like, a whole bunch on this, and I think it's great. I'm also super excited for having these other use cases, like scientific datasets and streaming or, like, sensor data or these other things in which floats would be pretty natural. I do I still I still have pretty strong, concerns about all this. I'm a little concerned. I mean, one is just, like, the JSON data model is just numbers, and they're not like, it's not the JSON spec that you don't put a dot if it's an integer. It's just undefined. Like, JSON doesn't say anything about that. And this relying on, like, well, we did much of testing, it turns out most of the time it mostly works. I'm uncomfortable about that. And I can kinda sharpen the the goal of this, like, mapping back and forth. I think in in some cases, you can think about, like, the CBOR representation that goes in the repository is kind of the canonical one. And as long as that's clean, everything else doesn't matter. Some of the cases in which it's important to be able to re like, round trip many times and continue to get the same output that we end up with in the ecosystem that might not apply in all use cases, but they do in some is you know, almost everything in App Proto starts as JSON. It gets posted to the PDF. So there's that JSON to p d JSON to CBOR that happens in the PDFs. Goes out on the fire hose, usually gets parsed back out into JSON. There's another time like, a time that it goes back is if the record was versioned, and you wanna check if it was the versioned. And this hasn't happened a ton in the network so far because not a lot of applications are supporting version, but there are some of these, like, profiles often can be updated in place. Anyways, you can end up with many versions of the same record, and you wanna know which version you have. And, basically, the way to do that is to turn it back into CBOR using a full Proto SDK. You put it back into CBOR. You hash it, and you get a SID, and you check. And some of the context in which you want to do that are like a moderation label, and you wanna see, like, does this label apply to this? Or you wanna say, oh, does this reply match back to this record post or this comment on a blog post or something like that? In those cases, this version checking is important. So, anyways, that's that's just to motivate. I I can feel a little like, we being pedantic about this? Does it matter that we can go back and forth? Because I do think it matters No.

[01:57:43] Dietrich Ayala: No. No. That we can do it. It definitely matters. The seaboard the JSON you get back from the seaboard Yeah. You can turn back into seaboard. You can keep going back and forth. It'll be the same. It might not be the the one you started with. Like, if you put in plus 10

[01:57:57] Brian Newbold: Yeah.

[01:57:58] Dietrich Ayala: Yeah. Plus ten and ten are both valid JSON. It's

[01:58:00] Brian Newbold: just one

[01:58:01] Dietrich Ayala: go to Seaboard 10, they come back 10 without the plus. They'll always be 10 without the plus.

[01:58:05] Brian Newbold: I think the confidence in that rests on we do everyone does everything correctly. Every every JSON implementation is doing this behavior, which isn't specified behavior. It's just most people do it maybe, and you end up with developers who are like, oh, my thing didn't work. I used this standard library in my obscure esoteric language, and it's not doing this thing. Whose fault is that? Do we go to the JSON implementation and be like, you're doing it wrong? It's not you know, they're just implementing the spec. I would also say this point at the end about the seabore and bytes parsing, that would apply also if we did some kind of syntaxial wrapper around floats. It's a equal point that, like, you already need to do this for bytes if you're taking the arbitrary app proto data as JSON and trying to parse it back out. You need to do this user land thing, this custom parsing out, and so we could be doing the same thing for floats as well. Sorry. That was, like, multiple comments.

[01:59:07] Daniel Holmgren: Hey. Daniel Holmgren from Blue Sky. I get, like, quite nervous about requiring people to have custom JSON parsers. Like that it it's just like a pretty I don't know. A lot of people are just gonna grab the off the shelf one. But everybody's already required to have a custom CBOR encoder. So I'm like, can you put more smarts in the CBOR encoding and less in the JSON parsing? And I'm wondering if, like, you know, like, you do these tricks of, okay, every float has to be, like, 64 bit float and every int is the smallest integer. Could you do something like if it's a whole number, if it's an integer, then it must be encoded as an integer. And if there's a decimal point, then it has to be encoded as a float. And that doesn't preserve, like, the fact that it's, like, necessarily always a float. Like, but but for any given number, it does preserve it round trip. But so it's like it I you wanna describe the field as being a float. You would describe it as a number, and that means that in the data model, it could either be an integer or a float. Does that make sense? I don't know if I'm being clear about that.

[02:00:11] Dietrich Ayala: Sorry. I don't understand the question.

[02:00:15] Daniel Holmgren: I'm I'm I don't know. May maybe this is better for the mailing list to to write it all up. But I basically wanna not have any smarts in the JSON parser and not have to have any smarts in the in the programming language. And if we can put all of this into the CBOR encoder, then I'm a lot more comfy and happy with it. That's fine.

[02:00:31] Dietrich Ayala: I mean, I think that's the proposal. I think the proposal is that the CBOR decoder or rather the yeah. I see what you mean. So right now, at the status quo is that if the JSON has a float in it, it just barfs. It rejects

[02:00:49] Daniel Holmgren: it. Mhmm.

[02:00:50] Dietrich Ayala: Could just make it a CBOR float. And the c board float

[02:00:54] Ben Goering: Base 64 loaded. No.

[02:00:56] Dietrich Ayala: What? Where where did base 64 come from? I don't know.

[02:01:00] Roman Danyliw: I don't understand that question either. Boring. No. No. No. No. No.

[02:01:05] Dietrich Ayala: Please no. Yeah. I'm I'm not sure this requires JSON parsers to be smarter or less standard. What I'm saying is that the JSON parsers already do this in almost every language or or can a pretty trivial thing to do in user space in the one language we found where you would need to. And, you know, if you click the the the GitHub link, it's a lot clearer what I mean. Like, every language has this distinction and most parsers automatically without even special configuration just accept floats in the JSON. Right? Like, the only thing that doesn't accept floats in the JSON is the lexicon tooling. It's like above and below, it's fine. This is what I'm gonna get.

[02:01:57] Daniel Holmgren: Okay. Sweet. Thank you.

[02:01:58] Shuping Peng: Okay. Cheers. So thank you. Thank you. And that is finished our session, and thank you all. Thank you. Cheers. A lot of fun this time. A lot of laughs. Yes.