Session Date/Time: 21 Jul 2026 12:00
[00:00:05] Job Snijders: Okay.
[00:00:12] Keyur Patel: By, by clock, it's time to begin. So let's do that. Whoever's in the back, would you close the door? Thank you. So as with all IATF sessions, we begin with the note well. These are your rights, privileges, and responsibilities. Make sure that you are aware of what they are before you contribute. Second, please make sure you treat each other well. It's okay to disagree with an idea. It is not okay to attack the person with the idea. Meeting tips. Everybody in the room here should have be using the on-site tool. We need this to manage the queue. We're gonna augment the note taking that Krishna has agreed to do with some of the things from the script and from the Etherpad. So let's make sure that you're using the queue management so we know who's speaking when. Off-site people, of course, need to use the full Meetecho in order to see what's going on here in the room. The agenda and slides are at the URL here. The notepad, which is used for the note taking, is at the next link. Please jump in there and help Krishna. If he if he spells your name wrong, fix it for him. If if he summarizes what you said wrong, fix that too. Alright. Zulip is the chat.
[00:02:01] Claudio Jeker: May I? The link are correct despite the title of the slides. Apologies.
[00:02:07] Keyur Patel: Oh, right. We forgot to update the title of the slide here. But the links are correct. Audio stream will be available here at this URL, and then there are the two links for the Meetecho session, the in the room site link in the full. So what do we get done since the last time? We've we've got a new RFC 9581 with the for the manifest number handling. We've got three documents in the RFC editor queue waiting for for them to get enough time to turn them into the final form RFC. We've got a document with the IESG now regarding the ASPA profile, which they hopefully will get to IETF Last Call and AD comments shortly. And we've got documents related to ASPA in working group last call now. So hopefully, those will be sorted out by the end of this week. Any questions on any of those? Okay. We have a very full agenda. So the agenda was posted to the mail list two weeks ago or so, and we had some discussion about that. And is there any agenda bashing before we get rolling through the agenda? Took two slides to put it up there. And that's then we get to the overflow things where we've got more requests than we, I am sure, have time. But it took three slides to put all of that. So we either need to be better about saying no to people who ask for slots or get more time. Time is over premium here at the IETf. As you can see, there's already a lot going on in each each session. Any I'm seeing no one get to the mic to bash the agenda. So we're gonna go ahead and get going with that. So the first the first one is the ASPA stuff. I think you're gonna talk about that. Did you get the new ones? Yeah. Okay. Yes. He's gonna give you control as soon as he gets it up.
[00:05:00] Job Snijders: Testing. One, two, three. So is this the correct version of the slides? Because I uploaded one minute ago.
[00:05:11] Alex Mendel: I accepted the slides.
[00:05:13] Keyur Patel: Check if if
[00:05:15] Claudio Jeker: you change it to anything.
[00:05:17] Job Snijders: Alright. We'll we'll see if it's the good version.
[00:05:21] Keyur Patel: And it's our intent. But that is Yeah.
[00:05:22] Job Snijders: Yeah. Yeah. Okay. So this slot is is sort of a open microphone session, I would say, to figure out, are there any outstanding issues that that we identify and and maybe can resolve immediately now, or are we good to go on publishing the ASPA cluster of documents? I'll I'll give a bit of an overview of where we are, what what the scope of the work is, progress reports, and then then it's open mic, for for anybody that has questions, comments, concerns, objections, approvals, tomatoes, whatnot. So SPA is supposed to be the world's premier anti route leak mechanism. We've been working on it for some time now. The original work started somewhere in 2018. As you can see, there was a little bit of a hiccup, in the progression, I think, mainly due to COVID changing the world a little bit. But here we are. There we're we're yeah. Years of work, and research have gone into this, and, we're, yeah, we're we're getting close to the end. So the current priority is to get a cluster of free documents published, specifically, a document that stipulates the ASN1, profile on how you publish an ASPA into the global repositories, a separate document on how you can use this cryptographic attestation in BGP best path selection, that is the ASPA verification documents, And then the third document is how do you transfer the validated ASPA payloads into the BGP routers so they can do the ASPA verification algorithm. These three documents are currently in working group last call. So this is a final opportunity for the working group to offer comments or suggestions or flag issues. We've not seen a lot of discussion or or support on the mailing list. So, if you are keeping notes on your to do items, I implore on all of you, add to your to do list to take a look at the ASPA documents and, yeah, tell the working group, what what your, thoughts are. Despite there not being a lot of feedback as of yet, what did happen, and for which I am super grateful for the chairs kick starting this, a few weeks ago, is, out out of working group reviews, materialized. So, GenArt, Routing Directorate, ArtArt, These are review teams that exist outside of the CIDR ops working group, and they took a look at, some of the ASPA documents and offered a a valuable feedback. So, like, one thing I found hilarious, is that a reviewer noted as, like, a a minor nits is like, instance of AS, you you, didn't capitalize the s. And then I was like, oh, that's so great because we are using customer ASes, that's CAS, and we have customer ASes abbreviated as CAS. So we're using the the abbreviation CAS for for two things, different things. So I ended up changing the ASPA profile, yesterday night to just fully write out customer ASes to avoid, ambiguity. And and this goes to show that that having outsiders review these documents is incredibly helpful because, yeah, I mean, all of us have been staring at this document for a while, and none of us were like, oh, that is And another reviewer pointed out, like, hey, the ASPA profile, is missing a proper security consideration section, so that has been added now as well. Then let's talk a little bit about implementations. In the SPA verification implementation section, a number of implementations, are mentioned. This means we we have the requirements of multiple implementations before RFC publication properly covered. I cannot speak for the the quality of these implementations. That is not my job. The the the goal of the implementation requirements in SIDROPS is to test the quality of the documents and whether multiple groups of people or AI these days, doesn't matter, can arrive at interoperable codes, and can read a specification and and produce something. So, there being five claims of implementation is is good news. And this also means that as this technology matures and the commercial, off the shelf vendors like Cisco and Juniper and and Nokia and Arista, and Huawei, when they take an interest in this technology that they have open source implementations that they can test against and compare whether their work is doing the same as, other implementations. So I'd say for SPA verification, we're we're in a pretty good spot, and that's that's good news. Then on the profile side, there's also really good news. There's multiple signers produced by multiple different organizations. Aaron, APNIC, and RIPE offer ASPA creation functionality, in a very easy to access place. The open source package. Krill has support for ASPA objects. There's a number of validators that have support for verifying, and decoding these objects, and there's many objects in the global RPKI today. And, again, this this shows that there's there's from different angles, different groups of people looked at the same specification and arrived at results that are interoperable. So here, I would say, cool. We we did it. Good work, everyone. Yeah. We got the old slides. Oh, no. I want the new slides. That's why they pay you the big bucks, man.
[00:12:24] Keyur Patel: Yeah. The big bucks.
[00:12:26] Job Snijders: You get to be first in line for coffee and lunch. Jeez. You just wave your chair, bitch. Right. That old slides. Would be super funny if I uploaded the old slides, and I'm now blaming you for showing the old slides.
[00:12:54] Keyur Patel: Oh, no.
[00:12:55] Claudio Jeker: That's not like it is the the v two, but uh-huh. I found the bottom. Let's see. Still there on the one.
[00:13:27] Keyur Patel: Yep. Same ones.
[00:13:33] Job Snijders: I think it's, like, v two at the end in the title.
[00:13:40] Claudio Jeker: Okay. It's this one. How do I move these outputs? Uh-huh. Maybe.
[00:13:53] Job Snijders: In Dutch, there's a saying, and that means that if you try something three times in a row, the third time you'll succeed. And it's yes. Yes. Yes. Alright. So, Tom, where are you? Raise your hands. Where? He's in the back. It's the the guy with the really nice haircuts and the glasses. He's trying to look like me. Tom delivered us a a pretty awesome thing. About five minutes before this meeting started, he finalized an implementation in interoperability report for, the RPKI to router, draft. You can check the latest version if you click the word here in the slides that you should be able to download from the data tracker, but it's hosted on the the SIDROPS Wiki on wiki.itf.org. And here is a a partial screenshot to to show you what you can expect. So what happens is that, in a virtualized environment, multiple different server and client implementations were launched and fed the same data. And then comparisons were made, like, is the same thing? Are they reacting the same way to the different events in the RPKI to router, protocol, finite state machine? So the long and short of this and forgive me if I'm mischaracterizing, the situation. If so, Tom, you need to come to the microphone. But I think with RTR, we're in a good spot, and there is sufficient interoperability among the the different implementations to proceed with publication. But there are some minute differences. So some implementations have not yet incorporated the very latest changes to the draft where two extra exits error reports were were added. Adding those for most implementers should should be very, very minor work, and it's on the exit path anyway. So, yeah, in that sense, it's, it's more of a nice to have than an absolute prerequisite to to proceed. But one minor thing came up where I think there there is, some con unclarity amongst implementers. And this is the case where if a client does a serial query to the server and the server returns a different serial, a different session ID than what the client had previously seen. In section five dot one, it says, you you do a a full, resets. The the the the cache must be flushed, you start a new session. And in five to three, a slightly different reading is that, the the server instructs the clients to to do a full sync. Now the ultimate outcome, whether you take the left path or the right path, is that a full synchronization happens. But in one case, you recycle the existing TCP session, and in the other case, there is a reconnect. And and this discrepancy, only came to light in in what was it? Less than twenty four hours ago. So I think this is the one point in the RTR draft where we need to figure out what is it we want, for to to be in the spec. But as I mentioned, the the the left or right, the the client and server arrive, at a synchronized state. It's just, the the number of steps between, towards arriving at that synchronized state, differs a little bit. So also for the chairs, this is the one issue that I'm aware of in in RTR, and I I think we can we should be able to hash it out this week. Yeah. And with that, we we arrive at, comments, and whatnot. I I know there are ASPA adjacent drafts existing in the data tracker, for instance, ASPA egress, ASPA slurm. And my my advice to the working group would be, let us finish these three documents, and then we after that is done, we can continue on the SPA adjacent, work. But we need, the core specification out of the door before we, take on more work. But, yeah, are there are there questions so far or or feedback? Applause.
[00:19:13] Claudio Jeker: Final change.
[00:19:14] Keyur Patel: Alright. One bug to fix.
[00:19:17] Job Snijders: One bug to fix is the long story short, and, please respond to the working group last call and and help the chairs understand if there's consensus to publish this as RFCs.
[00:19:29] Keyur Patel: So thank you. Are you still up? Yeah.
[00:19:41] Job Snijders: Oh, there's in the in the queue.
[00:19:44] Speaker 4: Hi. Just one question. The actual use of the ASPA in the wild seems to have resulted in some experience of, well, okay, unexpected behavior. And I that probably would be good to mention in the operations, operational considerations, section. Has anything that been done, in that area?
[00:20:35] Job Snijders: Yes. I I have good news to report. The ASPA profile in its introduction states that operators, if they create an ASPA object, must list all their transit providers. And the ASPA verification draft, mentions in the security considerations that if, if if there's a partial listing, then, some routes may be flagged as invalid. So both documents clearly express if you create an ASPA object, list all your transit providers, not just some of them, and that should mitigate, the the issue at hand.
[00:21:19] Speaker 4: Do we have any advice on how to deal with transit providers that are not really invited and thus are not listed and actually creating creating loss of connectivity as observed.
[00:21:43] Job Snijders: If if you are peering with a network and you believe they are leaking your routes, you have to shut down the session with the leaker. That is the solution that is used for twenty plus years. You disconnect leakers. But this was not a leaker. This was an an transit that was undocumented. So Well,
[00:22:06] Speaker 4: kind of Excuse me, boss.
[00:22:08] Maria Matejka: Skipping the queue. Job, please do not lie about what has happened. Oh. This was an undocumented transit which thought was transit and wasn't. Thank you on behalf of CCNIC.
[00:22:25] Job Snijders: There is difference of opinion on what certain documentation means, but the point is you have to list all your transits or disconnect from networks that are leaking your routes. That's all we can do. Alright. Next presentation.
[00:22:45] Claudio Jeker: There is another yeah.
[00:22:47] Speaker 6: Hey, Job. Sorry, like, I'll keep it quick. So, basically, with respect to the cache reset, right, Sriram, you mentioned that, like, the client is going to ask for a full update. So, like, we don't have currently, like, any, like, any threshold for which, like, max number of cache reset that the client can be entertaining. Because, like, if we keep on entertaining the cache reset, it might lead to overwhelming of the connection. Right? So do we intend to take this up to consider, like, at max, how many cache reset the client should be entertaining for setting the full reset to the cache server?
[00:23:25] Job Snijders: I I had trouble hearing the question. Did one of you catch the question?
[00:23:28] Claudio Jeker: Didn't catch it. Can you repeat, please? Speak up a
[00:23:32] Speaker 6: little bit. Yeah. Am I audible now? Is it better?
[00:23:36] Job Snijders: Or maybe you can type the question into the chats. Yeah. Did you hear the question, Jeff? Oh, Jeff, please relay the question.
[00:23:44] Claudio Jeker: Jeff is relaying the question. Thank you.
[00:23:45] Jeff Haas: Yes. On behalf of who I work with, his question is basically about rate limit impacts of cache resets. If you get driven into a situation where you're doing resets often, you know, the protocol's a bit ambiguous about the scenario and what each side should be doing about this sort of thing. Did I get that right?
[00:24:04] Speaker 6: Yes, sir. Absolutely. Sorry about my audio.
[00:24:09] Job Snijders: There are currently are no considerations, in in relationship to to rate limiting, expensive queries. And, I think this late in the game, it will be hard to come at a good outcome in a short period of time. So, I think this is a good topic to to explore in a a new Internet draft that I don't know. Rate limiting RTR RRTR?
[00:24:36] Jeff Haas: As two session two sentences that people move on. We did talk about this part of eighty two ten DIS because the conversation, for example, that you and I had had about what level of diff is, given server is supposed to be able to cache on and what happens if the server that's absorbing it as a client is running too slow, the protocol doesn't give enough hint hints about, what happens when you can get behind, how much chance you have to get behind, and then what chance to catch up, or to suggest to the server hold on to stuff because you have a slow client. So I agree with you. The the mechanism is clear in terms of what the protocol says, but as a development consideration, you know, this is worth working on later.
[00:25:20] Job Snijders: You for the question.
[00:25:21] Speaker 6: Thanks, Shaman. Happy to, like, join in in case you would like to take
[00:25:25] Alexander Azimov: a new draft.
[00:25:26] Job Snijders: Yep. Good to know.
[00:25:31] Alexander Azimov: You have to control.
[00:25:35] Job Snijders: Alright. Updates on an entirely different topic, a mechanical ex ASPA of the RPKI. How do we transfer those precious ROAs and SPAs as fast as possible to you? Quick recap on why the Eric synchronization project exists to begin with. We have rsync and RRDP at this point in time in the production environment. But there are some quirks with rsync. Rsync has a lot of states computation on the fly. So a client connects to the server, then client and server jointly spends a lot of IOPS and CPU cycles to figure out what is the minimum difference to be transferred from server to clients. So it's a little bit costly, compared to static, precomputed, content delivery. And more worrying, rsync has quite weak, consistency guarantees. Because just because a certain information object has a particular file name with a particular last modified timestamp with a particular file size, it doesn't mean that that, triple is represents the same content as in the past or in the future. So you can end up in situations where you try and synchronize through rsync, but there is no transfer of updated data because the algorithm, considers that there is no change despite there actually being a change. How can this happen? Imagine changing, for instance, the max length in a ROA or, in the same action, removing and adding a provider in an ASPA records, then it could be that the resulting DER encoding is of the same size as the previous version of the object. And if that's combined with, let's say, a timekeeping issue of sorts somewhere in the pipeline, you can end up with no propagation of information. Then in RRRDP, there there's a potential for retransmission of data. So for instance, an RP that had previously synchronized with RRRDP, RRRDP, it re fetches the snapshot, and it's a re fetch because the r in the cache, locally cached, information, the client already has most of the data that is in the RRRDP snapshot. And this is a result of the RRRDP enveloping, the the the way the the the snapshots and deltas are packaged. Don't offer the clients insights into what is in that delta I am about to download, and the clients can only learn what is in the delta until after they downloaded it. RRDP, does a full replication of the publication points history, but there's quite some churn in the RPKI. So this might mean that between the the previous time that you synchronized and the next time, there are events that have been overtaken by later events, and you end up downloading information that is just no longer relevant. R RRDP is tricky to combine with Rsync because of what I mentioned with the enveloping. So an RRDP doesn't exactly play nice with other means of transportation. And and this, to some degree, means that RRDP doesn't have the scaling properties I think we desire, because there's retransmissions, and and, the failure path is is a bit costly because you fall back to a full snapshot download instead of, something partial. Then both protocols have in common with each other that mirroring the the data without explicit coordination, is tricky. And what I mean with that is the publication point operator, they can choose to use, say, a CDN or a global anycasted approach, but clients cannot force the publication point to to bring the data closer to them. So it is not possible to do permissionless mirroring of the data in in the RRDP and and Rsync architecture. And, because of this lack of mirroring, not only can there be, latencies and congestions, but there's also some degree of fragility because there is no redundancy. It's a full mesh. All the clients need to talk to all the publication points. And if they cannot reach one publication point because there is a localized transient, connectivity problem, There is no other way to get to the same data. Now plan c or Eric synchronization is an approach to address some of the downsides that I mentioned in the previous slides. It's a very different architecture than RRDP, a very different than Rsync. It uses, Merkle trees to structure, and express to the client what the latest state is. And this way, the clients can very easily identify and very cheaply, this is important as well, did something change? And if yes, what
[00:31:13] Sriram Kotikalapudi: is
[00:31:13] Job Snijders: it that changed? And then they can fetch just the difference. And they can do this without establishing a session of sorts where there is shared state between the client and the server like there is with RRDP. And it's, it's it's static precomputed content. So in this sense, ERIC is a bit more like RRDP, where you can pre calculate everything and serve static files to the clients. So that is an appealing property. And it goes over HTTP, which is what most people like and prefer these days. So the deployment model is you have publication servers that use Rsync or RRDP. ERIC relays fetching the data over those transports, and they repackage the data in a way that it's suitable for, consumption by the relying parties relying parties. Sorry. IETf in China was a fantastic inspiration for this protocol because on paper, the ERIC protocol should be super performant and and exceedingly well. Like, it's you know, good job, Job. And here we were, Tom and I, hacking away at an Eric implementation, and we discovered that it was not performing as well as we thought it would perform. And I think this is because the the the previous meeting network had some some DPI component somewhere that, in effect, was a a layer seven rate limiter of sorts. Because what we discovered is that doing a 100,000 plus HTTP requests, is apparently too costly. And this was something that I did not really notice in Amsterdam on my high speed connection. So, what this taught us is that we need a way in the Eric protocol to make it so that with fewer HTTP requests, you can get more data in one go. And the overhead per HTTP request is significant. So if you fetch a two kilobyte object, nowadays, a TLS session establishment can consume, say, eight to 12 kilobytes. So you can see that there is an overhead of four to six times the actual payload that we're after. So it's really worthwhile to, give multiple objects to a single request, even if that means that you give the client objects that they already had, because it's offset by by the the request overheads.
[00:33:49] Claudio Jeker: Job less than two minutes and it Yeah. Yeah. So
[00:33:54] Job Snijders: what we made is something called, a write ahead log, which is, called segments in the ERIC protocol. Segments are an append only ASPA on the Relay. The clients can ask the Relay, what are what is your write ahead log from the last few hours? And then the clients can cherry pick which segments it wants to download to, prefill its cache. And segments are nothing more than concatenated, sequences of DER objects. An ERIC segment index is how you discover which segments exist. It's a very simple structure, and it's basically a sequence of time stamps and the latest index hash that, is associated with the latest writes to that write, to that write-ahead log. And here is an example, decoding. The ERIC index starts with a preamble about what FQDN is it about, what is the latest manifest this update contains in the segments. And then it contains a sequence of these tuples of the location of the segments and and the hash of the index associated with the latest write to that segment. It's it's static data to the in the in the sense that only the latest segment is changed. And every few minutes, the relay starts writing into a new segment, and that means the previous segments no longer change. Now implementation efforts are underway. There's multiple projects happening at this point in time, and I hope that we can report, to the working group, soonish with with some some full end to end implementations of the Relay and the client sites and invite more people to run measurements. And, once we are confident that this system has the properties we want it to have, then, maybe we can go for working group last call. But that's in a few months. Extra resources if you're interested in this topic, and apologies for overrunning by a few seconds.
[00:36:22] Claudio Jeker: Thank you. Thank you all.
[00:36:46] China Telecom Representative: Good afternoon, everyone. I'm from China Telecom, and my presentation is oh, sorry.
[00:37:02] Job Snijders: Yep.
[00:37:06] China Telecom Representative: Thank you. And my presentation is IPV six Mapping Prefix PDU for the IPK router protocol, and is the call answer.
[00:37:21] Keyur Patel: Okay.
[00:37:26] China Telecom Representative: What about this draft? This draft defines a new payload PDU that carries a set of IPv4 prefix authorized by single IPv6 mapping prefix. Its purpose is to carry MAP data from RPKI catch to routers in I p v six only environments. And why do we this now? Because and to IETF one hundred and twenty five meeting after after draft IETf MAP profile was presented to the working group last call. The chairs suggested that this PDO protocol extension should be delivered in parallel with MAP profile. So quick back up, there are two supporting working group drafts. What is v six ops framework, and another is IDR format six extension. So if you are interested about this, please examine why are data tracker, and let me focus on MAP mapping origin authorization. MAP is a side RP care object defined in draft IETF MAP profile. It also writes a p v six mapping fixed to originate mapping for one or more p v four prefix. The holder of p v four address block can authorize which a p v six mapping prefix is allowed to map that block. Why why does this matter? Because we without validation, attackers can hijack a p v four traffic by announcing fake mapping prefix. So here are PDO format, and it's a call and its key features. The PDO format follows the RFC eight to one area of structure, and the the key features include multiple IPV four prefix can be packaged into one PDU, and the flag field provides clear semantic for both announcements and withdraw of mapping authorization. So here are the PDU generation and processing. On the cache side, PDO is derived from validate MAP objects, and the cache generate one or more PDOs per MAP. And the IPV four prefix must be presenting in order. On the router side, when flags equals one, routers install mapping authorization in the local MAP database. And when the flags equals near low, routers will remove the corresponding authorization. For security for security consideration, the extension follows the RPKI and RPKI router protocol security models. Routers must be have a trust relationship with the catch, and the catch must validate MAP signature before generating PDU PDOs. So we also seek the need is clear, and we would like to call for adoption. And comments are welcome. We will refine it based on feedback. Thank you.
[00:41:37] Keyur Patel: Questions? Any concerns with the call for adoption? Okay.
[00:41:51] Jie Li: Yeah. Thank you. K. Hi. I'm Jeff from laboratory, and I will present our best validation with prioritized resource data. And now it is at version two, and it's a joint work across laboratories, Sunnet, Huawei, China Telecom, and Tinhua University. So here's a starting point that operators already mix in non apical data in their routing decisions. SLURM, which is RC8416, already lies operators for local assertions and filters into RPK validation. And operators may use local whitelist to avoid their to avoid black hole in their customers, and they may use our derived or third party fees where our ways don't yet exist. So the data is already here, but when operators may have some low confidence data and they want to monitor is in wide results. Once this data are folded into the slam generated via piece, such invalid results cannot be set be distinguished from that caused by the ROA data. So the gap is there is no shared semantic semantics. And today, each source is voted on ad hoc in vendor or operator specific ways, and there is no consistent rule for who wins when sources disagree. And most of all supplemental data can silently weaken RPKI with no one noticing. So this draft gives operators a common priority, a wide way to combine the feeds they already rely on. And this draft rests on three principles. We separate what operators decide and what's the framework guarantees. Firstly, the operators define the priority assignment. Operators defines the priority tiers and which sources sits in which tier. And secondly, the framework guarantees the priority ordering, which follows a higher priority result must not be silently overridden by lower priority data. And thirdly, the operator define the action per tier. Each tier maps to a local action, which is flexibility which is flexible on to the operators. And this draft in this draft, RPKS is authoritative, and the supplemental data must not silently override a validation result derived from the higher priority authoritative RPKI, and its default merge rule is highest priority ones where tiers disagree the conflict is a face to the operator rather than resolved silently, which means the lower priority valid cannot quietly rescue RPKI invalid. And here is an example, and we use three tiers and three different actions as example that operators may set the authoritative r r p k I as a high priority and easy and vital results will be dropped. And the operators may also use the IR draft feed and set with medium priority and deep refers invite results. And they may also have some experimental data and with low priority and they may only want to monitor is invite results. So the point is the flexibility that operators freely map each source to a tier and each tier to an action. The same inputs can yield very different locally correct policies. Okay. As for the deployment, we propose three models, and I want to free them as a cost ladder, not a rewrite. For example, the model a, it uses multiple caches and multiple tables. One validation table per priority with separate RTR sessions. Router differentiates via existing route policy, so it requires no RTR changes and no vendor code changes. And model b, routers pull per priority sessions and merge them into one table, And it also requires no RTR protocol change only with a priority aware VRPs. And model c needs a explicit metadata in RTR and with priority carried in band. Each requires RTR extension and the router and RAM parties changes. So baseline model a requires zero protocol and zero when the code changes. And we hear from Maria on the mailing list that Bert already supports this today. And for model b and model c, they are optional and can be future work. And this is how the three deployment models actually work. That model a keeps multiple tables on the routers and model b merges them into one table, and model c needs an RTL extension and carries per object priority in band. And this table presents exactly what each stakeholder actually has to touch. For model a, it only requires configurations, so operators can start with model a onto this stacks, bird or open BGP d, And model b is priority aware VRPs, and only model c touches the RTL protocol. And on the last Sunday's Hackathon, we build OpenBGPD prototype to realize these three deployment models, and we follow those three tiers of priority. And it is worth noting that in model c, we use the priority in Reserved byte of prefix PDU. Oh, and also models worked correctly, and this table can show a good example of how the lower priority evidence cannot override a higher priority value or invalid just as a third line that high priority invalid will cause the denied action. And although there is a medium valid action, the denied policy cannot over be overrated. And, also, as for the scale test, we run the full ROA and base 250,000 local supplements and our derived data separately. And here are two takeaways. As a left chart shows, the model a's configuration size grows to 16 megabytes with four hour a, while the model b and model c stays below one kilobyte. And secondly, as the right chart shows that as far as a full hour end to end validation time, which calculates from loading the input data to validating the routes on the routers, and validation times stay similar as the prioritized models have similar run times. So the semantics is richer, but there isn't significant additional validation overhead. And as far as outcomes, priority turns more data into differentiated action. And so now it is at version two. And after the meeting at Shenzhen, we have done three key changes. Firstly, we explicitly clarified that RBcast is authoritative, that this draft does not change global RPKS authorization semantics, and second is priority safe semantics that a lower priority data must not sell into your rate as authoritative RPKI. And, thirdly, is that we change the mandatory RTX extension to a model a b c later. Okay. And to sum up, I want to use three words. The first is necessary that operators already blend sources. We give it a shared RPK safe semantics. And the second is safe, that authoritative RPK is never silently overridden. And the third is a foldable. There's a baseline model a needs zero protocol changes, and model a b c is like incremental ladder. And the draft and our demo are available on the GitHub, and we would welcome the working group's questions, feedbacks, and discussions. Thank you.
[00:50:53] Claudio Jeker: Thanks. Job?
[00:50:59] Job Snijders: Job Snyder's. With the prioritization, I have some concern that inadvertently validation states are bleeding through into the routing data. Can you go back a few slides? There was, a little bit back. Thank you. A little bit back. Yeah. Sorry. Forward one again.
[00:51:21] Jie Li: This page? Oh, no.
[00:51:23] Job Snijders: This? Next one, please. Next one. Yeah. Uh-oh. This This this is the one. So the deep pref on IRR means that there is an implicit dependency on the existence of a ROA. And if the ROA is deleted, it causes the local preference in the network to change, which goes against the guidance in the recent draft, avoiding validation states in BGP that we approved as a to be published as an RFC. So yeah, I think this is one concern. Prioritization of the information can result in churn if ROAs come into existence or flip out of existence. And I think that is a problem with with the proposal at hand.
[00:52:21] Jie Li: Okay. I want to clarify that this is just an example. And maybe the operators may have their own local supplemental data. They come from maybe their customers or their friends with other SPs or some the AI info data and demand to use it flexible. And I think this framework is to give them this flexibility. Okay. Thank you. Thank you.
[00:52:47] Jeff Haas: Jeff has I will be brief. I think this is a real problem. We have customers that make use of static configuration to override RPKI. I agree that we need a taxonomy about how we should be able to do that. I agree that information inside of RPKI router is helpful. So I support the draft. I think some of the contents we'll want to talk about, but I think this is good work to move forward.
[00:53:13] Jie Li: Okay. Thank you. Thank you.
[00:53:17] Claudio Jeker: Okay. Thank you.
[00:53:18] Alex Mendel: Let's move forward.
[00:53:32] Tomoki Yoshikawa: Okay. Can you hear me?
[00:53:36] Claudio Jeker: Yes.
[00:53:36] Tomoki Yoshikawa: Hello. Yeah. Okay. So we're started. Hello, everyone. My name is Tomoki Yoshikawa from University. And today, will present my work on post quantum signatures for the RPKI. Since version zero one was submitted, the member with discussion has changed the direction of this work. So I will present both the results of the current software and the proposal for how to organize the next documents. And next page. Okay. This slide shows the current status of the draft. Budget zero zero was published on June 19 and the current revision version zero one was published on July 1. Today, I'm not asking the working group to accept all parts of zero one as a single specification. Instead, I will use zero one as a source of the current code and measurement results and then ask how the work should be spread into separate documents. Here's a prompt for this talk. I will first present the proposed document spread and then explain the main points that should guide the choice of post-quantum signatures for the RPKI. Yes. This slide shows my proposal to split the current work into three documents. The first would be an informational document on PQC considerations for ZARRpki, hovering requirements, possible approaches, measurements, and operations without selecting a mandatory next state. The second document would describe mixed certification change as a migration method that is not tied to one specific algorithm. It will explain how an RPKI certification path can change algorithm at the CA boundary, and it upgrades the migration model defined in RFC six nine one six. The SAR document will eventually define a specific RPKI signature suite as standard struct work after the base specifications, software support, and the needs for real use already. So these are proposed scopes rather than three complete drafts. So the first question for the working group is whether the spread is useful. And for the rest of this talk, I will mainly discuss materials that could go into the FAST document. The main idea is to study each approach against the specific needs of the RPKI rather than starting with the winner chosen in advance. The study should cover pure post quantum signatures, composite signatures, and related approaches such as a neural scheme. It should consider both security and results from real implementations because it possibly grows RP verification time, supporting signing systems and HSMs, and tests between different implementations can all affect the final choice. The purpose of this document would be to support Rated standards work, but not to select the mandatory next step by itself. Yeah. Yeah. Okay. And this this slide explains the reason for this work on this call I'm considering. NIST has standardized post quantum signature algorithms and the pure ML DSA profiles for Exod five zero nine and CMS are now being drafted. The base standards are zero for moving forwards, while the prod RPKS data relies on RSA 2,048 with SHA 256. The main question is how to study and introduce new signature methods without changing the parts of the RPKI that do not need to change. And this will cover signatures in LISO certificates, CRLs, certification request, signed object, and a BGPsec router certificate. It does not cover BGPsec upgrade signatures and it does not change RPKs, CRLs, RRRDP, Rsyn, RPK payrolls, or router behavior. And this slide shows the main experimental case used in version 0.1. Version 0.1 used composite ML DSA 65 with ECD SAP two fifty six as a difference point for testing and comparison. Both parts of the composite signature must pass for edition and for protection against false signatures, the design aims to remain secure as long as at least one part remains secure. It also allows one RPK object to carry one composite signature instead of publishing a cross cow version and the separate postcumin version of the same object. There are however two important remits and first the pure ML DSA profiles for X. Five zero nine and CMS are now already RFCs while the composite signatures used in this experiment are still internal drafts. And second, an existing RP that supports only RSA cannot validate the composite object. For this reason, I think this profile is useful as an experimental reference point, but I am not proposing it as a mandatory next seat today. And this table as a composite result to my pure RSA and a pure MLDSA measurements. I will focus on the highlighted MLDSA 65 plus p 256 case, which signed 100,000 message in about forty one seconds and verified them in about eleven seconds. And I I received verification to cover one second in the same test, so the cost on the RP side is an important point. The model repository size for the highlighted case is about four times the RSA baseline, although this is not a measurement of the complete RPKI repository. The point of this table is not to declare one row of winner but to show the positive posterior growth and repeated RRP verification need to be part of the selection process. And the figure shows why the migration method should be considered separately from the choice of signature algorithm. The parent CA remains on the current RSA suite and assigns a transition certificate using RSA, while the child CA public key inside that certificate use a new composite suite. An updated RP first by this transition certificate with the parent's RSA key and then use the composite public key to validate the child's CA and its object. The algorithm, therefore, change out the CA boundary, which allows different subtree to move to the new suite at different times while only one version of each subtree is published. The trade off is that an old RP cannot validate the migrated subtree, so RP support must be ready before CE switch its subtree. This message is not tied to their DSA or composite signatures and could also be used with other current and next suites. That is why I propose moving this design to a separate Mixed Certification chain storeraft. The figure and the main idea are based on Derek's work. And this slide shows what I have completed and what still need to be done. The published call now cover draft 19 compose signature creation, validation, and performance test but it does not yet provide for our PKI object processing. The next step is to use composite signature in the RX. Five zero nine certificate, CRLs and CMS signed objects followed by full validation test from CA to RP. At the same time, the current document should just should be spread into clearer scopes and and the studies to cover a wider set of signature approaches. This right is my question for the working group. The most urgent question is whether the proposed three documents bridges the right direction. I'd also like to know whether any important requirements or selection points are missing and which signature profiles oriented approach you need more study. And I will welcome comments on any of these points. And that concludes my presentation. Thank you to everyone for your review version zero one. I'm joined in everybody's discussion and comments, implementation feedback, and the call also is very welcome. Thank you.
[01:01:43] Keyur Patel: So RFC six four eight nine talks about how to do algorithm transitions with RPKI CAs. And I think that you need to restructure this to align with the process described in that document.
[01:02:01] Tomoki Yoshikawa: Yes.
[01:02:17] Bob Beck: Bob Beck, OpenSSL. I'm not a routing person, but I know slightly enough to be dangerous. When you're go back to your little slide there where you were doing the the transition from the the RSA CA. Back back
[01:02:31] Speaker 6: to RSA.
[01:02:31] Bob Beck: Yeah. Yeah. There we go. Okay.
[01:02:33] Tomoki Yoshikawa: Yes. Okay.
[01:02:33] Bob Beck: So you're aware here there is, of course, a downgrade problem.
[01:02:38] Tomoki Yoshikawa: Yes. Yes. Yes. It's a big big problem. But,
[01:02:43] Bob Beck: basically downgrade problem, I mean, if if I can if if RSA gets broken, like, the other thing doesn't matter because I can put my own MLDSA in there, and I'm happy, and you'll trust it.
[01:02:59] Tomoki Yoshikawa: Yes. If we adopt composite signature, if RSA has is broken,
[01:03:10] Bob Beck: we do use No. No. No. No. No. I'm not I'm not talking about composite versus versus noncomposite. Pretend it doesn't matter. I'm talking about that RSA certificate on the left that's pure RSA that you are using to trust that transition certificate. It has signed over that transition certificate. So if RSA is broken, that's great. You've got a perfectly secure post quantum algorithm in there, but you'll believe any key I give you because I can sign whatever I want.
[01:03:37] Keyur Patel: Bob, I don't think the arrows are certificates.
[01:03:41] Bob Beck: I thought this
[01:03:42] Tomoki Yoshikawa: was a trust chain.
[01:03:43] Bob Beck: Oh, okay. This is not a trust chain. That's my my my misunderstanding. Because one part of your draft does say that you should not accept current suite and next suite for the same node in the tree to try to avoid the downgrade problem at one node in the tree. My concern there was that that's good, so I I won't have a ROA misspecified on a downgrade. I can't hijack something there. But if other parts of the tree still accept this, I could have this row appear anywhere. So nothing's really safe until the entire tree does not accept this. Correct?
[01:04:25] Tomoki Yoshikawa: Yes. Yeah.
[01:04:28] Keyur Patel: Thank you.
[01:04:34] Tim Bruijnzeels: Okay. Tim, responding to Ross, I guess. Yeah. So mixed trees or separate trees because of the existing document for algorithm rollovers talks about having separate trees, a new tree where you do new things. I think that warrants a much longer discussion. We should probably take it to the list. But if you go down that route, there are quite a number of things that we need to touch in other spaces.
[01:05:05] Keyur Patel: Correct. I agree with you too, and that's why I wanna have the discussion on the list. You can taste this. You're gonna have to click this thing first. Now he's working. Okay.
[01:05:30] Jie Li: Okay. Hi. Jeff from laboratory, and and this is ASPA-based AS_PATH Verification for BGP export. And it is now at version five, and this is joint work across laboratory, Tsinghua University, c z dot ni NIST, and Huawei. And let me start with a quick recap. Our idea is simple. That is to apply the ASPA based ASPA verification to outbound e b g p updates. The existing working group draft ASPA-based AS_PATH Verification verifies at ingress when a route is received, and this draft as a check at egress on the outbound AS pass just before export to a neighbor AS. And why would we verify at egress? Here are two motivations. The first is a partial deployment. In real world, some ingress routers don't yet support as per verification or b two p rows, while some egress routers may already do. So the egress verification then covers roles that would otherwise leave the AS with no ASPA based export check. And the second is operational assurance. A route valid on ingress can turn invalid once exported. That may be because of the local misconfiguration AS migration or just omissions or mistakes in published as per records. And such errors might be subtle that may appear only on roads sent to one neighbor. And I like origin errors. They seldom surface in public BGP analysis tools, so the egress verification is where we can catch those catch them. As for how it works, the egress verifications sees after the local ASP prepending and before propagate to the specific e b g p neighbor, it verifies the exact AIs pass the neighbor will receive. And the verification choose the algorithm based on the relationships. It matches what's ingress as part would apply, how the route being received in the reverse direction. And the relation could come from BGP capabilities as per objects or local configurations. If the two AS have a complex relationship, we can sub segregate the session into regular rows or decide perfect fix, or operators may use the downstream algorithm to avoid false positives. And I want to make it very clear about the position of this egress draft. So it changes no ASPA semantics and disables nothing. It complements ingress ASPA. And if the OTC attribute is present, it must not be overridden by egress verification. So this draft acts as a local safety night at export. We can say that egress ASPA is to ask for verification. What's RFC 9319 that RPKI ROV at export is to the ingress ROV. K. So what's new in this version five? We've done three key changes. The first is we repositioned the neighbor AS-augmented verification, which I will discuss later. And the second is we expanded the optimizations. And third, we refined the motivation and the whole structure. As for the first change, neighbor AS-augmented verification is to prevent the neighboring AS number to the pass, which helps detect a missing provider AS in ASPA records. As shown in this table, which we analysis before that, if there is a ASPA omission, which means one AS missing one of its provider is as per records, such omission could be caught one or two hops away with existing verification. But with neighbor AS-augmented verification, it can be caught locally, and this is where the operators can actually fix it instead of two hops downstream away. And this mode most naturally used in the tight offline monitoring setups, in-line egress enforcement stays local policy choice. And the second change is about making egress verification scale. We propose three approach approaches to optimize egress verification. The first is centralized or detached, and the the work can be offload to a route reflector or central node, or it can be fully detached via BMP only notifications. And the second is deployment controls with enable, disable, changing enforcement mode as deployment involves. And the third is partial verification. We can reuse the ingress verified prefix or verify only the local part with only full verification during AS merges, list, or remembering. And egress verification is deployable today onto independent open source stacks, BIRD and OpenBGPD. Egress verification is already available on BIRD that as per check can be caught anywhere in the export filter, any pass including a prepended one. And, also, we implemented a demo on the open p g p d that reuse its ingress verification and as a outbound call set. And it supports directional wear, and it gives OTC precedence, and it can reevaluate the as part it can reevaluate the road when there is as part updates. Also, it supports three modes of monitor or enforce. So it is now at version five, which includes all the planning, revisions across, the previous IATF meetings, and we wonder why the the design considerations are complete. Are there any cases, risks, or interactions we've missed? And we humbly welcome the working group's perspective of this draft. Thank you.
[01:11:46] Keyur Patel: Alexander.
[01:11:57] Alexander Azimov: Hi. Thank you for the presentation. I was so I have some concerns about this document. It's not about the general idea, but some parts of the wording make me any. Especially on the draft current version says that the egress ice path verification should be performed on all roads possible illegible for being propagated at egress. I wonder what is possible illegible for being propagated. It is not specified in the document, and I'm not familiar with this step. It's it's not in it is not in the slide. It's it's in the draft.
[01:12:42] Jie Li: I know. I know. Oh, okay. And I I will check
[01:12:47] Alexander Azimov: the So the draft uses term possible illegible for being propagated, and it's making me cautious.
[01:12:57] Jie Li: Okay. Okay. As for the details of the draft, I will check it later carefully, and we we can discuss
[01:13:06] Alexander Azimov: Okay. And can you also get back, I think, to
[01:13:10] Job Snijders: a few
[01:13:10] Alexander Azimov: a few slides back, please? When you were discussing how to retrieve the local role. Yes.
[01:13:25] Jie Li: Oh, this this right.
[01:13:27] Claudio Jeker: Yeah.
[01:13:29] Alexander Azimov: I'm getting I'm not I'm confused to this with the second point using ASP objects of local AIS and neighbor AIS. There was a reason why such technique was not suggested in a SP verification document. I think it should not be here too. So it's okay to use local role. It's okay to use some configuration. Using a space objects for this purpose, I believe, is unsafe.
[01:14:05] Jie Li: Okay. Thank you.
[01:14:12] Jeff Haas: Jeff, as Alexander said about half of why we wanted to say, this slide is good. The consideration about local AS is very important. You know, this is not well considered in the ASPA verification for import. The other case that should be discussed is remove private. So we're deleting private AS numbers.
[01:14:35] Speaker 6: Okay. Thank you.
[01:14:41] Alexander Azimov: Okay. Okay.
[01:14:42] Jie Li: Thank you. Thank you.
[01:15:03] Sriram Kotikalapudi: I'm Sridam from NIST. This talk is about so so we have ASPA, which is already very well positioned, well thought through. The algorithm works well. So so nothing changes about ASPA. I think Jia also tried to point out that with the egress verification also, nothing changes about the ASPA verification as we have it today. So that's something that we want to continue to, progress and and complete, as we are doing currently. So once ASPA is available as the base, then we can think about some enhancements that may fill some gaps in ASPA. And the the gap that I'm talking about here is not something that we haven't already acknowledged in the ASPA verification draft. We have acknowledged this. And here, we are contemplating, okay, very ASPA already moves the needle significantly in terms of detecting route leaks and a variety of path manipulations. But in in addition to that, there is one little gap that that I will describe in a moment. So we want to be able to see if we can have some solution for to fill in that gap as well. So you'll see the details of that in a minute. So as I said, the the ASPA based AS path verification can detect all route leaks. It can also detect what we call forced origin and forced path segment hijacks. Those are well described in in the ASPA verification draft. So it can defend against those as well when the update is received from a customer or a lateral peer. That's when we are doing the upstream verification. When we are doing the downstream verification, there is a possibility that as I mean, per alone cannot detect these types of hijacks, when it is received from a provider of to a customer. So these are essentially fake link or force peering attacks, which I'll show in a moment. So our goal is to detect and mitigate these fake link attacks received from any direction from a provider, customer, or lateral peer. So as per route leak detection capability, the one goal I mean, one clear goal here is to keep the ASPA route leak detection capability intact. We are we are not going to meddle with that. That is the base that is important to all of this. So we have acknowledged this problem of of of one little gap with with ASPA verification. We try to find the up up ramp on the left and find the down ramp on the right. And these are with the help of ASPAs, customer to provider authentications, ASPA is used. At the top, the these two ramps may merge in onto one AS. In that case, there's no issue. If these two ramps merge on two different ASS, like three and four, ASPI of as three may have ASPI or three may not have asp-, it doesn't matter. In either case, the following happens. Three, a test that, that four is not a provider or three says that three has no ASPI, in which case we don't know whether four is a provider, customer, or lateral peer. In all those cases, the if if four is in the path, then no doubt, that that this path is valid. So only question is, is four faking a connection to three? Is four really connected to three? So so that's the issue. So far out, somewhere else, there is a x, a s x, and there is a a s y, which is a customer. One, two, three are another part of the Internet. And if x somewhere else can fake a link to three, so one has ASPA, two has ASPA, three may or may not have ASPA. So x can send it down to y and the ASPA as it is as it stands now would y would y y would determine that it is a it is a valid path. Because three to four can be customer provider or lateral peer, it doesn't matter. I'm sorry. Three to x. And in all those cases, y would determine the paths to be valid. So there is a fake link issue here. X can fake a link with three or x can fake a link with one or two, and it can get away with it in the downstream direction, but not upstream. That is well recognized. And also, maybe four is not somewhere far away. Four is in the path. And, however, four received a longer path from one to two to three down to four. And four now fakes a link with two. By doing that, it shortens the path and sent it sends it to its customer. Again, in it's a downward direction attack. And the the other alternative that the customer has through six is valid, but it is not able to differentiate. It in fact, it it she's a she's sees a shorter path from four, although it is invalid and accepts it. So the and then a customer can also take advantage of a non adopting ASPA. In this case, six can take advantage of seven not adapting it. And so then seven prioritizes the customer route over another route it may have received from elsewhere and and sends that to five. So five is is all five is indirectly attacked by six. Although, it does it still does come through seven. The important thing here is partial adoption. So we are we are in partial adoption, these things can happen. And we are this example serves to show show that if you have something else like what we are proposing, ASRA, then even in partial adoption of ASPA, if you also have ASRA. So if one in this example has both asp- and ASRA, this can be detected by five. So the solution is in the form of a new RPKI object, which we call autonomous system relationship at his authorization. It allows the registration of customers and lateral peers. So either the ASPA or the ASRA must confirm a BGP peering link with the next AS in each hop in the AS path in the forward direction, the direction of the path. For example, ASI to ASI plus one, we will, call it a a fake link, if ASI's ASPA and ASRA both exist and neither of them attest to a relationship with a s I plus one. So that's that's the key. That is the central point of the solution. ASPA alone method guarantee its guarantee is that the ASPA is feasible. ASPA plus ASRA method may gives you the additional guarantee that if there is a fake link somewhere in this path, then that that can be detected. So we have four subcategories of ASRS. ASRS c, LP, CLP, and TICS. C is for customers, l to to register the customers. LP is to register all lateral peers. And as for CLP, in case the subject AS, is not willing to disclose the customers, separately as a list, they can combine lateral peers and customers into one list and not be explicit about who their customers are. They don't want to reveal in this in that case, they can register them as a combined list. We won't be able to differentiate them, in the algorithm, which is a light bit of a disadvantage, but that can be option for the for the AS. And finally, ASRA TICS. TICS is trans transparent IXRS Rx. And the the proposal is for the TICS to declare itself as a TICS using as ASRA, and then it can optionally register all its RS clients. So by doing that, the following happens. If two and three are clients of a TICS RSA s 100, What the TICS RSA s does is to register a TICS in mentioning itself, 100, in it, and it can optionally include all its RS clients. It does not have to. It can optionally include them. In that and then ASPA two and as a s two and a s three, those are the RS clients. They are appearing at this, at RSAs. They create ASPAs and include 100, in in the ASPA, in their ASPA. So we know from the TICS as ASRA sorry. Yeah. In the TICS ASRA, we know that from that, we know that, 100 is a TICS RSAs. And by looking at the ASPAs of two and three, we know that they are both pointing to 100 TICS RS, so they must be, clients and therefore they are effectively lateral PS. So right there, we know that there is a connection between two and three, albeit through r a r s a s, but they are lateral peers in effect. So where this helps is that we we so this clearly, like, I like briefly describe the solution to you. There are detailed algorithms and I have a bunch of slides in the backup, which we'll not go through today. The algorithms are well described there as well as in the drafts. So the algorithm basically is able to catch, these fake links based on those based on the use of both ASPAs and ASRAs. So one other side benefit of ASPA plus ASRA is that combined, they detect together more route leaks than ASPA alone in partial deployment. So this is partial deployment example, for example. And in this, one has ASPA, it may not have ASPA, that that would also be fine. Two has aspa and ASRA, that is the key. And if two two happen to leak it to three, and then three passes it down to, so three itself is not doing aspa, but it but four is doing it. Four would not be able to detect it based on aspa alone. But if we have two, both two has both aspa and asra, then that confirms and that confirms that three is a lateral peer. For example, two's ASPA says three is not a provider. And two's ASRA says three is the lateral peer, either through an Rx, TICS RS AS or or directly. We we can catch both from ASRA. So in that case, we clearly see that there's a lateral peer and this path should be invalid and Ford would be able to detect the invalid in this case. So there there are benefits for early adopters. You just need the ASPA the ASPA SRAC adoption by two ASS. And between those these two ASS, if other ASS don't adopt, we still would be able to detect fake links if any happen. And I mean, the the adopting AS from there if there's a fake link that that can be detected. Thank you and I'll be happy to take questions.
[01:27:18] Keyur Patel: This
[01:27:23] Maria Matejka: is Maria from BIRD. I'd like to I'd like to note that I have been turned down with the request to prevent single AS NES numbers in the ASPA verification. And now we are proposing a mechanism which is going to put into the routers potentially much bigger and larger and more complex databases where I would I would have to look whether the feasible whether the feasible parts of the a s path are beating here or here or here or there. And if any one of these fits, then there is probably valid. I'm seeing a lot of problems with this. I am not against working this way, but I am a little bit concerned that this is going to have a lot of friction.
[01:28:15] Sriram Kotikalapudi: You are what? So I'm sorry. The last sentence? Okay.
[01:28:20] Maria Matejka: It's gonna generate a lot of friction.
[01:28:24] Sriram Kotikalapudi: Okay. Thank you.
[01:28:27] Tim Bruijnzeels: Yang Fei from laboratory. ASPA is conflict when ASPA conflicts with ASRA, what's the results?
[01:28:40] Sriram Kotikalapudi: Okay. Good question. One thing very clearly we state in the draft and make it absolutely clear that as ASR would never take any precedence over ASPA. So and also, we use ASR only in the forward direction because the we are looking for possibility of ASI plus one faking regarding its relationship with ASI. So ASI plus one's ASR, we we we never consider that. We only consider ASIs, the preceding ASS, ASPI and ASRA. So thanks for that question. We we took care of that in the draft.
[01:29:21] Jeff Haas: Jeff Haas, I'm here to say that I think that you're gonna want to allow the conflicts using slide 10 as your example topology, just to comment on this, a complaint I've had for a very long time about ASPA is Gau Rexford is a business practice. It has nothing to do with correctness of the BGP protocol. The consequence is anytime you have a complex style relationship that violates Gau Rexford, which three and four would have if they are, you know, not ASPA peers in terms of the Gau Rexford, but are valid BGP relationships, ASR can model this. You, therefore, if you want to endow that use case and allow for the complex behavior or gradient ability at a per destination basis to override that ASPA should not be enforced for a given destination for a given ASR path. We can take up the case later, but, you know, I'm just raising it that
[01:30:10] Claudio Jeker: way.
[01:30:10] Sriram Kotikalapudi: Yeah. I would like to understand the details of that. I but I think we did discuss complex cases fairly well in the draft, but maybe it still requires some more thinking, and I'll be happy to discuss offline.
[01:30:24] Thomas Schmidt: Thomas Schmidt. I wonder from a trust perspective, if I look at this declaration of a lateral peer, wouldn't it be necessary to have this declared by both? Because a lateral peer is equal, I mean, equal relation. And if two ex declares it, but three says no, then it actually is not valid. Right?
[01:30:46] Sriram Kotikalapudi: Yeah. We are already trusting two's ASPA. So we take two's ASPA into consideration and we trust it fully. So so two, if two also has ASRA and it declares three to be a lateral peer or or no connection with three, then we want to trust two about that more than we would trust three because three is potentially the the one who is faking a link, not two. So we are look we are going in the direction of the flow of the of
[01:31:16] Thomas Schmidt: the update. Example, free simply has no SPA. It is not faking in this in your picture here. Right? So I I was just wondering I mean, declaring an an upstream provider is is is a once one a s statement and lateral p p, I would expect it to be a two a s statement.
[01:31:34] Sriram Kotikalapudi: Yeah. SPA may not adopt one situation. And if it did adopt and if it it confirms that two is a lateral p, then that is even better. But we we don't explicitly need that information for the verification.
[01:31:50] Alexander Azimov: Alexander. Two comments. First of all, I admit that there is a significant progress by introducing this special object for the transparent access. At the same time, from other operational side, it makes things significant different from the ASP. Because in ASP, customer is responsible for keeping its ASPA object up to date. Here happens a delegation. It's who should keep the object up to date for for its customers. And if something happens, if if it's not up to date, it's will be kind of indirect responsibility. So it's kind of delegation, and it's very important to point out. Second comment is that you are well correct while describing the limitations of ASPE, especially for prefixes that are received from providers. But it's also important to to highlight this ASRA doesn't fully cover the problem because it stay in a window for replay text. Not replay. It's not replaying, but to to construct a path that will be valid even in ASRA world, but it will be just a cons artificial construction. And also getting back to IXs, it doesn't cover the scenario of selective period inside the IX.
[01:33:25] Sriram Kotikalapudi: Yeah. I think I take the comments well. So we can I think you and I need
[01:33:31] Bob Beck: to talk talk a little
[01:33:32] Sriram Kotikalapudi: bit more and discuss about the IX cases, which I think you are concerned about? In this one, I don't know if you are concerned about, the RS revealing all its, clients. It doesn't have to. That's optional. As long as, two and three attest that the 100 is is is their RS, They are good. I mean, we are good. And it so therefore, it doesn't have to be a full adoption on part of all the clients of an RS. Some of them adopt. That that is fine. So again, like, the partial deployment benefit is that ASR might be might be adopted in just a handful of ASS, and they benefit for sure immediately. So it doesn't have
[01:34:16] Claudio Jeker: to Maybe Sorry. Maybe you continue offline. Yeah. We are still running very late. Sorry.
[01:34:44] Ying Su: Hello, everyone. I'm Ying Su from NCGC lab. Today, I will introduce over new draft requirements for API relying party. This draft aims to refresh RFC 8897. RFC 8897 provides a single reference point for API API software requirements. It helps implementers identify RP relevant requirements across multiple specifications quickly. From a from the experience in developing our own RP software, we found such an entry point is very useful when building RP software architecture. However,
[01:35:25] Jie Li: I have to
[01:35:25] Ying Su: say eighty nine seven was published in 2020. Since then, the API ecosystem has evolved significantly. Several new specifications and updates have been have introduced a new requirements for RP implementations. This include several new signed object types updates to the synchronization mechanism, validation procedure updates, RTR relevant work, and the new operational considerations. So the IP requirement landscape today is different from the one described in RFC eight 8897. The goal of this draft is not to redefine RP behavior or create another layer of RP requirements. It is to update the RP oriented reference map prior to the by Office eight 8897. It organize RP relevant specifications around the RP processing model, including repository synchronization, certificate and processing, sign object validation, value to the cache distribution, and the new operational considerations. So the authoritative requirements remain in the individual's verifications themselves. The overall structure of this job will still follow RFC eighty nine seven. Most most existing areas are returned and update, while some areas have a major updates due to the changes of the API standard. For example, API repository synchronization and the signed object has has a major have the major updates due to the new repository synchronization mechanism and the new and object type. We also add we also add two new subsection in the introduction. One describing the changes since the RFC eighty nine seven. The the other one clarifying references to the active working group working group drafts. We also add a new area which called operational and the management requirements for RPs. This slide summarize the major technical updates. First, the API repository synchronization incorporates updates relative to the RRDP, such as same origin policy, desynchronization recovery, and the new synchronization mechanism, Eric. Second, the trust anchor has involved the new signed object, the TAK, and trust the anchor selection mechanism. Third, certificate and the processing has been up updated according to the path validation updates and the style handling work. For signed object, this draft also updates the requirements for the raw and manifest and adds the new signed object type such as as part k k and RSC. This draft also updates the look the valid cache distribution mechanism will be the RTR version 2 and add a new section operational consideration. We also we also include some some active set of working group documents. These documents are included because they are directly relevant to to the IP implementation, but these documents may we we may consider them provisional. If the content or status of this document changes, the the corresponding reference may be update, replaced, or removed. Today, the modern IP software is a continuously operating security component in the production network. So operators may need variability into a repository synchronization, validation procedure, and the validate out out of both state. So this draft also include three operational considerations. The first one is validate validate the cache export. The second one is observability and the diagnostics. A third one is audit trail. This capability help operators compare results, troubleshooting, and analyze RP behavior. Finally, this draft is not intended to replace the underlying specifications that RP developer developers must read. It is it is to help them identify which specifications are directly relevant to the IP implementation and helps them understand how this how this specification fit together from RP workflow pros perspectives. Oh, yeah. So we we would like we would like to gather some feedback feedback from the working group on this document. Thank you.
[01:41:01] Keyur Patel: Job?
[01:41:06] Job Snijders: Job Snyder's. I will reiterate some feedback I shared on the mailing list. In the development of an RP, the development team needs to make an overview of what RFCs exist and which ones are relevant and which ones are canceled out by other RFCs. And this this is a fair bit of work. I I know from experience, and I think you you know from your own experience, this this is a fact. It it takes a little bit of research what to implement to make an RP. But the specific RFC of RP requirements, literally was outdated already before it got published. And I think that in, trying to update the the that RC, we will repeat that that dynamic of it's always being yesterday's news. So I do appreciate the work that has gone into, like, what does a modern RP look like? And I think that type of helicopter overview is is valuable. But I think RFCs are not the best publication venue because of the slowness of RFC publications. So I think another way to to use this work and help future developers could be to use the the CIDR ops wiki and have a page that outlines like, this is the current state of the art for an RP. These are or optimizations you can think about or tricks like, yeah, let's jointly document how our our collective RP implementation experience. But, yeah, specific on on updating the existing RPRC requirements documentation, I I think that that RC should just be marked historic, and that would help, developers more than updating it every five years. So but this is my take. And, thank you for the presentation.
[01:43:09] Ying Su: Oh, thank you for your comment.
[01:43:12] Claudio Jeker: Okay. Thank you.
[01:43:29] Alex Mendel: Hi, everyone. My name is Alex Mendel from TU Dresden, and what I'm about to present is joint work with Thijs from RIPE NCC. So the research question that we kind of set out to answer was how fast does the data plane react to RPKI VRP changes? So effectively, you do this by having a trace route before the VRP change, and then a couple of trace routes during and after the VRP change to figure out how the paths changed. Now, this is not entirely new. There's prior work that asks the same question, and that work usually uses RPKI beacons, so prefixes they control, and they sign the ROARs for themselves. And they flip it, valid, invalid, and then run trace routes to that prefix from their vantage points. But that limits the topological coverage to the upstreams of the AS that is running the BGP, the RPKI beacon. Now it would, of course, be better to use in the wild prefixes because then you cover a more diverse topology, but this introduces uncontrolled behavior in that we do not know when these prefixes might flip in their valid validity status or the announcements flip in their validity status. So the question is, how can we start measuring on the data plane before the RPKI ecosystem sort of reacts to a VRP change? And our answer to this is effectively run a very optimized validator, and instead of looking at the public repository, look at the CDN origin of the RIPE NCC report and get every news a minute early compared to the rest of the Internet and the API ecosystem. And so by looking at that origin with the optimized validator, we then run trace routes from three globally distributed vantage points to the impacted prefix of a change for one hour every three minutes. Now, you may ask correctly, so does that work? And the answer is yes, at least for our three month observation period. We took a list of public validators, so run by Cloudflare, by RIPE, and a couple more from some IXs, and we compared when we observed a change versus when those validators publish or push a change, and we were at least a minute faster, and in the median, we were five minutes earlier than those public validators. So we good enough. Now what did we observe? First of all, we ran our measurements for three months from the November 2025 until the twenty first January 31, and we, in the RIPE repository, observed 25,000 ish VRP transitions that affected VRPs. So we only are interested in changes that changed the VRP. We don't care about CA resigns or valid until changes. We only care for those that can impact the validity in that moment. And of these 25,000 transitions, we have 1,493 transitions that go from or into an invalid state and that do not have another covering announcement. So these prefixes should be disconnected or connected to the Internet if they flip into or out of invalid. And, yeah, we measured how fast they appear and disappear. So this is from an invalid state into a non invalid state, so either valid or not found. And the graphs are separated into upstream count or number of upstreams, and this is for the AS originating the prefix that is going valid. We derive this number of upstreams from RIPE RIS data, so we may be overestimating or probably overestimating the number of single upstream networks, but it is good enough for this comparison. And for single home networks, announcing, again, the prefix that has changed its validity status, the median time is twelve minutes until 50% of our vantage points are able to reach that prefix. And for the multi homed case, the median time goes to nine minutes, so you're quicker if you have multiple upstreams, although if the originator has multiple upstreams. And we can look at the same thing in the other direction. So going from a non invalid state to a invalid, the times increase quite drastically For if the announcing AS is single homed, the median time until 50% of our nodes are only able to reach the prefix as thirty minutes. And for multihome systems, this time increases to forty two minutes. Also note, we only ran the trace routes for an hour. So the big blob at the multi multiple upstream's sixty minute is not because everybody withdraws at that point, but because I stopped measuring at that point. Another fun thing that this data has is that I can look at how fast large networks withdraw invalid routes or sort of stop routing them. And this graph shows the last time a probe that had observed one of these networks in its trace routes stops seeing that network in its trace routes when going from a not invalid state to an invalid state. And we can see on the left the fifteen to twenty minute range is probably what we expect networks running ten minute polling intervals on the validators to look like because there is a bit of head start that our method has, and then there is sort of, yeah, somewhere around the ten minute reload time. And then the other providers are probably running less often, pulling it off. Now what do these results mean for you? One thing is that deploying RPKI or setting a ROA during a hijack might not be quick enough. So, for example, the sort of well known Route 53 hijack a couple of years back only lasted for roughly two hours. And in our data, after forty two minutes, 50% of the vantage points were still able to reach these invalid prefixes. But there's also an easy solution to this, create ROAs before there is a hijack happening so you don't have to rush this much. And secondly, have validators re fetch in shorter intervals, so ten minutes or less is probably a great call. Or your RPs, so this is what the validators want.
[01:51:03] Keyur Patel: Okay. Thank you. Thank you.
[01:51:04] Claudio Jeker: Jeff, please, one question.
[01:51:10] Jeff Haas: Jeff has one question, slide 12. Yes. So we're seeing using Arelion as the example and Cogent as well that the stretch on this is very long. You know? Yep. I would expect this to be a tight blob. Have you formulated any theories why that Yes.
[01:51:26] Alex Mendel: The blobs for the ones that are slower is actually they're benefiting from the quicker networks. So if one of the quick networks is right in front of the announcing AS, announcing the invalid, and they quickly turn off the announcement, and that is the last announcement for this prefix, then all the others get the same quick timing because they get a withdrawn route and also withdraw the route. So, essentially, if if we have a if if we have cogent NTT, the origin is announcing invalid, and NTT quickly withdraws that announcement following the invalidation. This quick withdrawal also happens for Cogent, and that's why they have a blob at the same area that the quicker ones have exist or exist. Yeah.
[01:52:17] Jeff Haas: Okay. To summarize, thing number one, the closer to the core you are, the better you're gonna have results. The second problem is this is a last man standing problem. BGP normally has that, and this just happens to be showing that the last man standing problem is actually far worse Yes. Than they should.
[01:52:35] Alex Mendel: That is is that yeah. Yeah. Yeah. So the that that's quite correct. The this was pretty clear that this would look somewhat like this, both with the this being fast, this being slower because that's just what BGP does. It looks for connectivity. So establishing connectivity is quick. Losing it takes a bit. RPKI just multiple, like, increases this in a lot of magnitude.
[01:53:00] Claudio Jeker: A quick one, John.
[01:53:03] Job Snijders: I I just wanna say thank you for this research. I think it is really awesome how you managed to convey the dynamics of the system. So thank you so much for coming here and sharing this study with us. Thank you.
[01:53:18] Claudio Jeker: Thank you for being.
[01:53:19] Keyur Patel: Thank you. Yes. We have seven minutes left, so we're gonna try and squeeze one more presentation in. Sorry if there isn't time at the end for questions. We'll see.
[01:53:33] Speaker 6: Hey. Am I audible? Hi, everyone. So I see the oldest slide to be uploaded. No problem. I'll just brief it out for everyone. Like, the it's basically RPKI over the quick protocol, for my, like, coauthors. I will be going through this draft and my implementation. Okay. So before we dive in, deeper as to how, like, we would like to help in our AI, a couple of beneficial points that we see is, like, Quick has the inbuilt security mechanism, the TLS encryption for which, like, unlike TCP, like, where we need to have an enhanced, like, separate support for security, like, TCPO or MD5. We don't need any such extra, like, security mechanism here. Second is basically the restoration and establishment time that is pretty fast. Like, we don't have the blackout time over here, really. Like, when the service are getting restored, like, we have the validation done pretty fast. The next one would be, like, here, we maintain couple of streams within one connection, which means that we have control on the flow of packets for each of the streams and we we have a proper control of them. And quick connection ID allows session survival when the client's network address changes. So these are few of the benefits that we can take a look for QUIC. Coming to the establishment now, the thing is, like, we need to share particular identity, a RTR-o-QUIC, to the cache that we want to support RPK over quick. So, basically, the router should be acting as a client while the cache must be the server. Along with that, like so there is we we don't get the early data implementations over his like, we should not be allowing that, like, without proper setup. We should not be sending any application data towards the cache. Similarly, for the termination, we follow the, like, RFC for quick as defined in RFC 9,000. And it's and for the idle time out, we definitely maintain that a pretty high value so that, like, if we do not receive any packets for a given time, we do not terminate the session. Okay. So couple of PDUs, we can distinguish them between like, we can majorly broaden the two categories as data packets or control packets. And ideally, like, we can punch everything in a bidirectional channel as I already mentioned. Like, there will be couple of streams within one connection. It can be bidirectional. It can be unidirectional. We can punch in all the, like, PDUs in a bidirectional stream if we want. But the benefit or, like, for better efficiency, we can use unidirectional channel for the data packets because it's gonna flow from, like, just the crash server towards the router. Whereas the control channels or, like, packets the control packets will be going by the control channel, which needs to be bidirectional and need to be brought up soon after the connection establishment. Right? So there are two ways for data channel, which we can operate in. So we can either spin unidirectional channels soon, like, before, like, we send the cache response from the cache server, or, like, we can spin up bidirectional channels accordingly. Like, that we will take a look in the following slides. Alright. So for if we if we go for unit unidirectional channels for data packets, ideally, that should be spent before we send a cache response period towards the router from the cache. So after we get a re request for a serial or a reset query, while we send the cache response, we should be sure of spinning the unidirectional channels for the data packets.
[01:57:55] Claudio Jeker: Right? And Oh, you have one minute left if you want to go maybe.
[01:58:00] Speaker 6: And and before, like so that that's that's basically to to fasten up. Like, I'll just keep elaborating on it. So we can go for the unidirectional channel or, like, it can be bidirectional channels. For bidirectional channels, like, ideally, we spin the bidirectional channel before sending the serial of the reset query. This termination happens once like, the session can be terminated after receiving this end end of data PDU from the cache server, or it can be kept as per requirement. So four data channels for four different types of data packets, v four v six ROAs or out of the PDU or ASP APU. So, yeah, I did a fast pretty fast, like, collaboration of it. So the main thing is that we should be sending a flag, our TROQIQ, towards cash to resemble that we want to support.
[01:58:56] Claudio Jeker: I'm sorry. I apologize, but the time is over. I would suggest next time to go on the why, not on the how we do the quick thing. You ask for five minutes, you have five minutes. I mean, not 12 slides. I'm sorry. I'm sorry. So we take it to the list to continue the discussion on on quick. Okay? And time is over.
[01:59:23] Job Snijders: So should
[01:59:24] Claudio Jeker: meet in San Francisco. Thank you for all.
[01:59:29] Keyur Patel: So thank you all. We have less than a minute left between the whole session. So thank you very much, and have a good rest of your meeting.