Session Date/Time: 19 Jul 2026 16:00
[00:00:08] Chairperson: Ladies, gentlemen, nonbinaries, all of you, grab your sheets because we're about to get started. This is the first hot RFC where every slot we had available has been filled up. So if you've been here before, you know that one of my great pleasures in life is keeping you strictly to four minutes, and I expect the audience will be helping me with that. First off, for many of you, this is your first ITF meeting. Welcome. We are very glad that you are here. You will be seeing this slide frequently. And if you are not, the people running the meeting have done something wrong. This is our note well. This is basically the high level description of the process and policies that you are following by participating at the IETF. Please read through it. Get familiar with it. It does matter. The QR code will take you to it so that you can actually get more information if you are unclear on any of the aspects thereof. If you're here to present because you've got this really cool idea that you want to see taken through the process to become an RFC, the IAB is ready, willing and able to help you do that. They will have a new work help desk this week. It's a great place to ask someone who's had to do this far more times than they would like to admit on Monday and Thursday. Swing by. It's downstairs.
[00:01:35] Yuning Zhang: It's by registration.
[00:01:36] Chairperson: By the registration desk. And again, they would be more than happy to talk to you. There are some ground rules. I've already touched on one of them, but let's go through from the top. This is basically a request for conversation. Right? This is a moment for you to say, hey, I've got an idea and to get some advice on where to go with it. Some people have come here because one of the dispatch groups said, you what, you should take this to Hot RFC. Some people have come here because, hey, they actually have a new idea and they're just trying to get a sounding board for what you think. Right? You get four minutes. If you hear something really interesting and you want to ask them about it, remember who they are because you don't get Q and A during Hot RFC. Find them later. Buy them a coffee or a tea or a beverage and I'm sure you'll have a great conversation. We do have all the things are four minutes. Four. At the end, there will be a please applaud slide and when you hear that, if you are speaking, it is time to stop. Let's practice.
[00:02:40] James Kunle Olurandari: Well
[00:02:45] Chairperson: done. With that, let's go ahead and get the party started with Dan York.
[00:02:55] Dan York: Hi. I'm Dan York. I work on this on the staff of the Internet Society, and I wanna talk oops. Thank you. I wanna talk about the Open Fibre Data Standard. So it begins with the question of how can we build a resilient and reliable Internet if we don't know where the physical cables actually are. We have wonderful maps of submarine cables. Thank you very much to telegeography. We've got lots out there. But when you look at what's happening on land, it's not that great. There aren't really pictures out there. There's lots of maps that are out there in various different forms, in various different ways, but they aren't actually in forms that we can use. Many of them are maybe interactive maps on a website or they're JPEG or they're PDF or they are KML, which is the format from Google Earth, or they might be Excel sheets. There's all sorts of things that are out there in different ways. So the challenge is that we don't really know what the infrastructure is like out there in ways. So back in 2022, a group of organizations started to go around with creating something that they called ultimately the Open Fiber data standard, which is a language for describing the points of presence and the spans, the interconnections that go through there. So this is OFDS. It's at ofds.info. This out there and you can go there, read the spec, look at the pieces, see some of the tools that are being used. The it describes the nodes, the spans, the house there. It also has additional language around is a line leased versus owned, you know, ways that you can go and describe this. It it is a JSON file. It's also a geo package, some pieces like that. The people looking to use it have been regulators who are trying to understand the resilience of their network in their space in some way. Companies that are looking to buy connectivity, and they want to be sure that the route they're buying from ISP a is actually physically distinct from the route of ISP two and not just the same ones leasing the same physical lines. Funders who are looking to go fund new connectivity, wanna be sure they're paying for something new in some way. Researchers, people preparing for disasters. Look at something like, you know, if there's a wildfire or a flood that's gonna wipe out, you know, forecasts to come and wipe out a piece of highway, what's that gonna affect? What's the outage? What's the piece of there? So the folks who have been involved with OFDS for that past couple years are actually here at this ITF meeting looking to try to understand, is there a fit with the ITF to bring this standard into here, to bring the document, the specification, and have it go through the formal SDO, the formal standardization process, which in the case of here would be, you know, mailing list, boss, however many something working group, etcetera, etcetera. So are there enough people here interested to start bringing this into that process? Now wait. You might say this is physical layer. We don't do a lot. We don't, but we have. We brought GeoJSON through here. Anybody who's been around with GeoPriv and the eCrate work that we did back in the February, that was a lot of this kind of work that's there. So we have a side meeting on Wednesday, February. Actually, in here. Right? This is Grand Climb Paltry. Next door. Next door. Next door. Two on at 02:45 on Wednesday. And that's it. You can find contact me, Dan York. I'm York like New York. York@isoc.org or Steve Song. Some of you may know he is song@isoc.org. And that's all.
[00:06:27] Chairperson: Excellent. Awesome. Alright. Shengmen Jing, BGP Secure Origin Delegation Authorization. Authorization is one of my favorite words. This is gonna be great.
[00:06:41] Shengmen Jing: Hello, everyone. I'm Forest from Library. Today, we introduce BGP Secure Origin Delegation Authorization (SODA). I have seen 1930 recommend one perfect one perfect has one or has one only one Origin AI. But in today's Internet, let it make multiple origin AS is common. So the challenge is how to distinguish the legitimate mouse from hijacking. The current ROA has deployed, but the measurement says that more than 10 more than 700 routing companies per day were class classified are invalid, which is higher than the actual number. In other case, most cases are false positive. There is a sim simple example. Two ACs in in a commercial partnership announced the same prefix, but only one AS is authorized by our way. So the another the announcement from another AS were classified as invalid. Our is a standard object data states which AS is a solid add of two orange eight eight IP free prefix. In some real world, the seniors, the authorized AS needed to collaborate with other AS to achieve business considerations. But the current API cannot explicitly express express its authorization. So we intro introduce the BDP suda, a new BDP path attributed to distinguish legitimate most from route hijacking. The workflow is straightforward. First of all, the delegator generate and assign the sudo token using the router router key, and then the delegator attach a sudo attribute to BDP update message. When the receiver receive the BDP update message, it perform perform a standard Away and check and verify Suda attribute. That's all. Thank you. Welcome your comments, feedback, and collaboration. Thank you.
[00:09:37] Chairperson: Great. Thank you. Alright. Maic Sans Fleishier, Beyond Throughput Switching Efficiency for AI Data Center Networks.
[00:09:55] Maic Sans Fleishier: So hello. I am Maic Sans Fleishier from Université Grenoble Alpes, and I am working on BGP as well. So Where? Okay. Thank you. Okay. It takes time. Sorry. So you know that we have root leaks and path hijacks in the Internet. So we still have some issues with BGP. And there's no mechanism for ASPAS validation in the protocol. There's, like, very good work on ASPA being done, and it's getting maybe sometimes sooner adopted. But we think that ASPA does not properly handle complex relationships, which are very which are intricate relationships between IS outside of just a verifier provider and the client. And so the answer to that, I mean, we propose like path validation, PAVA. So what the idea is, is that the received AS path is cut into segments. So here, like, a s six, a s six will receive the path 65721. And we cut it in segments of free with for each AS in the path. And then send that segment about the prefix to each of those ASes in the path. And those ASes can give an answer about whether they consider this little segment valid or not. This is done through the DNS query. And so the validating AS here, AS six will ask all of those and get a number of answers from each of those a s's. So for a s two, it will be seven to one for the prefix that was advertised 1.1 in slash 16. And after getting the answers, we have a verification algorithm that will process the answers. And this brings the advantage for the ASCs to only maintain local information in their DNS system. And so the queries will happen through DNS. There's much more details in the draft. And so I will do a little presentation at the working group. We'll be also at the other BGP rated working group. And so if you have, like, any interest in in this or to talk about it, so come talk to me or send me emails at my address. So I'm welcoming contributions, talks about it, and if you have any idea for implementations or stuff like that. Thank you.
[00:13:02] Chairperson: Well done. Alright.
[00:13:09] Nengye Ye: Good evening. I'm Nengye Ye from Shanghai Jiao Tong University. I will introduce a speeching efficiency framework for AI data center networks. In AI data center spotting LLM training, both communication algorithm and network capture can be shaved network efficiency. First, consider the already primitive. In a ring based implementation, a GPU may receive intermediates or redundant data. This data discarded during reduction. This creates networks and endpoint activity that does not directly become useful computational inputs. With in network computing, reduction can be performed in network So the GPU receive only the final reduced result. These two implementations can therefore exhibit very different efficiency characteristic. The same issue appears at the network architecture level or three d Torus can be well suited to adjacent communication, but may be less suitable for imbalanced workloads. A real optimal architecture can better accommodate imbalanced workloads while potentially incurring additional multi hop forwarding. These differences may be not visible from CherryFoil or utilization alone. Our question is, can we define a metric to quantify network efficiency for LM training? Fireworks and bandwidth utilization are useful starting point, but they show how active the network is, not how much of their activities post computation or whether the deployed project and capacity match the world. The first challenge concern computation. As shown on the left, the bytes carried by collective communication do not necessarily become input to the net neural network computation. In all reduce, for example, and reduce intermediate data consume network resources by its not retained afterward. This is why we should distinguish computational effect data from other traffic. The second challenge concerns world project fleet. As shown on the right, DP, PP, and TP produced traffic that is often sparse and uneven. Static top topology properties alone cannot tell us whether the topology fits that pattern. What matter what matters is how well the deployed network support the actual training traffic. This tool gets motivated proposed switching efficiency framework. Switching efficiency, Eta, is the computationally effect data in equipment per unit time divided by total switching capacity. The left side, define the numerator. D is the full gradient or activation tensor per GPU, and delta d I is the data equipment from communication primitive I that is retained for subsequent computation. It's well defers across communication primitives. Intermediates or and reduced data may consume network resources but does not contribute to Delta D I unless it is retained. The right side defines the denominator. Some of RPs at the theoretical equals rate of the included switching resources from SPI and LIF switches to in service switching. It therefore measure effect effect data delivered per units of provision switching capacity rather than generic network utilization. Switching efficiency can be decomposed into three factors. Its overall value showed the conversion from provision capacity to effective data, while the factors show where the loss occurs. To calculate the factors, we measure four quantities. The ratio between these four quantities yield the three factors. Port utilization as how much aggregate capacity was used. Routing efficiency as how much forwarding work was required for data received at endpoint. Data efficiency as how much received data was retained as effect data. A low point utilization, point to idle, or poorly distributed speech in capacity. A low routing efficiency can reflect tour with transmission loops or topology.
[00:17:21] Mateo: So
[00:17:22] Chairperson: the slides are always available on the data tracker. So for further information, go there. Alright. This time, I'm actually gonna read the correct ones. Next speaker is remote, I believe. Mateo, we have you.
[00:17:37] Mateo: Yeah. Hello. I'm Mateo Karék, French cybersecurity students. And today, I'm going to present you the DIVE protocol. Okay. Do
[00:17:53] Chairperson: you have control of the slides?
[00:17:54] Mohit: No. I can't. Okay.
[00:17:55] Mateo: Oh, I have control.
[00:17:56] Chairperson: No. We alright. We have control. Tell us when to change slides.
[00:18:00] Mateo: Okay. Next slide, please. Thank you. So today, we have software update agents, package managers, and IoT devices. And what is crucial about this, to manage it, we have updates. But updates are critical moments because an attacker could compromise an update artifact, and this is why you have today multiple protections, of course, TLS. But we also need to check if the origin is compromised or not. This is why we have embedded keys and CA based signings, but we need a public key on the agent in the client. And it adds complexity of curation, and the key could be compromised, which is also critical. And it can also add external cost, external dependencies, and complexity. And this is why small manufacturers deploy an update on this. Next slide, please. This is why I want to add a new concept that uses the DNS to store the public keys. Next step, please. So we have the origin servers that delivers the artifacts with a signature and the key identifier directly in the headers. The dev client will then get the public key from the DNS and see if everything matches and accept our rejects versus next status.
[00:19:36] Mohit: Next. So
[00:19:38] Mateo: we have two types of records. We have the type records that is basically the configuration and that says, hello. I'm the protocol. I exist on the server. And the key records sub contents, the key there could be multiple key records. And DNSSEC is mandatory, of course, because without DNSSEC, an attacker could compromise, all the chain. Next slide, please. So this is, so Dive has a lot of flexibility features. The first one is, of course, key rotation. To rotate the key, we add a new key. This is where we have key IDs. Then we sign all the artifacts with a new key, and we remove the old key, and we have successfully retreated our key. We have also we can restrict the key to a specific subdomain. This can be useful if we have a big organization and we want to add specific keys for specific areas. And we have, of course, a report only mode. This is for the first migration to avoid the then your service attack. If you didn't configure Dive correctly, you have only reports and the resources must blocked. Of course, you need to remove it for production. And the keys must never be stored on the origin server. And, of course, that doesn't depend of on big adoptions. It's something that you add independently on your own company or infrastructure for your local infrastructure or IoT devices. Next slide, please. So if you are interested in this concept, you have here all the links to the Internet draft and open source clients and a dedicated website, which also supports the protocol. And I'm looking for feedback and eventually web protocol. Thank you very much for your attention.
[00:21:48] Chairperson: Fantastic. Andrzej Duda, self certifying identity and capabilities delegation.
[00:21:56] Andrzej Duda: Thanks. Hi. I'm Andrzej Duda, and this is a work with my colleagues from Technical University in Grenoble. The principle is that the principle is that we want to provide identities for AI agents. The way we define them, it's a hash of a public key and the representation that will go that can be stored as a name DNS name. The property of this identity is that an agent can prove the possession of this identity because of the possession of the secret key. Here is the big picture of of the whole framework. So imagine that we have several agents, AI agents, and, there is one which is a scan agent. The example is taken from, network security. So imagine I can have a scan agent that discovers the detect agent and wants to establish secure communication and delegate some execution task to the scan to this detect agent. The providers of the agents and all entities in the in the system will have identities, and their relationship between the identity and the public key will be stored in this trustful mutable store. The first implementation of the of such a store can be DNS with the DNSSEC. Another implementation may be based on IPFS. So how to set up this secure communication? So the idea is to have the first negotiation between the client provider and the server provider, and they will establish a capability that will be given to agents. And the client agent will provide this capability to the execution agent. And everything can be verified, the identities of the providers, identities of the agents, and set up finally the the secure communication. So this is the the two protocols that I I combine in this this framework. Here is an example of the how we can store the identities with some metadata information and public keys. Another way of storing this information would be with the IPFS and IMS. It's a little bit more complex. It requires two records that will be stored in d a in g h GHT of IPFS and the binding between the public key and the identities based on IAPNS. So to finish, this kind of architecture provides a good way to representing identities and coupling public keys with identities. The relationship between identities and the keys will be stored in the trust for mutable store. And we can provide some trust between entities like client and server providers based on the fact that they have identities, names, and they can trust each other based on on names. And finally, we can provide the protocols for establishing the secure communication for executing tasks. Thank you very much.
[00:25:57] Chairperson: Alright. Andrew Fregly, PQC DNSSEC.
[00:26:02] Andrew Fregly: Hello, everyone. I'm Andrew Fregly here to present on PQC DNSSEC MTL update. So a thirty second review on the post quantum photography challenges facing DNS. In the near future, quantum computers are expected to break the RSA and ECC algorithms used in DNSSEC today. And as a result, many governments and industry participants are targeting 2029 through early twenty thirties to migrate to PQC solutions to be prepared for such a future. Now standardizing into deploying solutions takes time, often years. So for PQC DNSF, you really need to start working on solutions now, and work has been ongoing. But PQC algorithms today have unique challenges for DNS. Some algorithms have large keys or large signatures that exceed the DNS over UDP traffic profiles that we have become accustomed to, and some algorithms take longer to sign in or verify than RSA or ECC. As a result, there is no one PQC algorithm that behaves like RSA or ECC without one of these challenges. Now as I mentioned, work has been ongoing on PQC DNSSEC at the IETF primarily through the PQC DNSSEC side meeting. And out of the IETF one twenty three side meeting came this strategy document that proposes a routine performance and resilient fallback options for PQC DNSSEC. At this time, we do not have any recommendations for routine performance. Those algorithms are still going through standard standardization and finalization, but there are resilient fallback options in the form of MLDSA and SLHDSA. Now we as a community wanna standardize both, at least one from PQC routine performance and one resilient fallback for DNSSEC purposes, especially to meet the cryptographic agility objective in DNSSEC, which is critical because that enables signers to switch to a different algorithm if their current one is compromised or in the future, of course, if a new PQC algorithm is standardized that behaves more like the routine performance options. Now in the meantime, we can reduce the impact of some of the PQC performance penalties by applying techniques like the Merkle tree ladder approach that we have been proposing, which uses a Merkle tree structure to amortize the cost of large signatures across a message set, for example, a zone file. So work on MTL has been ongoing at IETF as well. Since IETF one twenty four, we've created a new MTL specification for DNSSEC that is more self contained than previous iterations, and it uses a MLDSA, which aligns well with many of the other efforts ongoing for PQC at IETF. We've also provided new DNS specific options and details around how to use EDNS zero to tie in when to accept or return full and connect signatures, one of the ways that we amortize the cost of PQC algorithms, and we describe how MTL can reduce the overhead for both static and dynamic signers, especially when multiple RR sets are in the same response. This rep does have IPR. It's public. It's royalty free, and you can click on this link to learn more. So why am I here today at HARRC? Well, our next steps for MTL is that we really want to get community feedback. We're in the process right now of updating our open source code to align with our latest draft, and any feedback we get at this I IETF, we can then incorporate and include in the next round of open source updates we make. We will also be looking for adoption of our draft by the IETF working group responsible for p q c DNSSEC efforts in the near future. And importantly, we want to continue to support community efforts. We've been partnering with open source resolver implementations in academic institutions to test both the MTL concept and PQC DNSSEC concepts more broadly. And if you're interested in MTL or PQC DNSSEC, we encourage you to reach out. We're always looking for more partners and more folks to test with. We will be at the hack demo happy hour tomorrow and at the PQC DNSSEC side meeting on Thursday, or you can reach out to us over email attached to the draft. Thanks for your time.
[00:29:59] Chairperson: Alright. Oh, look. You just popped right up. There you go. This is great.
[00:30:05] Mohit: Hello, everyone. I'm Mohit. I'm from NITK Suratkal, India, and I'll be talking about interactions between the L4S and FQ based queue disciplines. This is a very new work. We haven't done any activity on this, but we are looking forward to, conduct some experiments. The reason I'm here at HotRFC is to tell people that this is something that we found in the last ITF as something that would be interesting to explore and there are a bunch of experimental tools that are available but we haven't seen any real deployments. So, first of all, we have seen a widespread deployment of the flow queuing based mechanisms such as FQ codel and CAKE, common applications kept enhanced. Most of these Linux distributions they support both FQ codel and CAKE and they are not just there but they are actually being used in several ISP networks and also in home gateways. Cake is particularly very popularly deployed in home gateways. We also have FQPY, which is an active Internet draft in the TSVWG, and currently, we are seeing its adoption happening in some of the Wi Fi networks in India. To summarize, FQ based Q disciplines are getting deployed. On the other hand, the deployment of L4S has also gained a lot of attention. So we have deployments of L4S by Comcast. We have often seen the numbers that Jason has presented in the TSBWG working group as to several million homes in US now support L4S. We have Apple, Charter, T Mobile, Nvidia, all of them are working on L4S supported by major vendors like Ericsson, Nokia, Samsung, and also built into modern DOCSIS 3.1 and slash 4.0 equipments. So vendors are also supporting it. But then the core question that lies in front of us is that when both of these are seeing active deployments, how can we bring them together, and how can L4S be integrated into FQ based queue disciplines? That's the most important question that we wish to answer. The integration is not seamless as it may seem like, because both of these approaches have their own architectural goals. FQ's goal is to provide fairness by isolating traffic into per flow queues, whereas L4S goal is to achieve ultra low latency by keeping the buffers shallow. And there are also some nuances in the scheduling logic. In L4S, you have a dual queue scheduler. It has an L4S queue and a classic queue. The scheduler prioritizes the L4S traffic, but the coupled ECN marking pushes back the L4S traffic and somehow it balances what L4S traffic does along with the classic flows. But that's not true with FQ scheduler. FQ uses a deficit round robin scheduler and it's a slightly different modified version. It's not a plain DRR. If a flow is a sparse flow or a new flow, it gets more priority than a old flow. Then the question arises that if L4S gets into one of the FQQs, how will it be treated? We do not have answers for that. We do not know how per flow management of the queue will hurt the L4S performance or will it give any kind of benefit. We also do not have any information of how congestion marking will happen, which is definitely decoupled in this case. We don't know about the sender pacing, and we also don't know how deficit round robin will behave for new and old flows, particularly when L4S is involved. So there's a need for experimentation in this, there's a lack of practical deployments, although FQ codel in Linux kernel supports L4S marking, it just tells you how to enqueue an L4S packet but it doesn't tell you how it will be dequeued. So there is no practical information on that. We are looking for support in terms of fine grained experiments using L4S with FQ based queue disciplines, and also some case studies from real deployments would be really helpful. If you are interested, please get in touch with us. Thank you.
[00:34:13] Chairperson: Great. Santosh.
[00:34:21] Santosh: Good evening, everyone. I'm gonna talk about the AI inference streaming, the protocol gap. Sorry. So the protocol gap. So why do we need a standardization here? And, you know, who will benefit from that? It is really about how the LLM servers stream the tokens to the client as well as the ID back ends to the front end. Like, you know, what what kind of protocol they use, and why is it important for us to standardize it, and who benefits from this. So currently, if you look at the AI traffic that is from the server to the client, you'll see that the transport layer is all standardized. That is HTTP two, then you have TLS, and then SSC. These are all standardized. But what happens above that is something very unique to the or it is proprietary to that particular servers or how the vendors really want to, you know, have their own schemas, have their own protocols aboard. So what we observed is that some of the vendors want to keep a particular stream for a long term, whereas, you know, others keep it for the short term. Also, we observed that, you know, some of the vendors have started using Connect plus Proto. That is a protobuf with the Connect protocol. The issue there is that it doesn't work with h t p two sorry, h t p one dot one. So this is the protocol breakdown. OpenAI, Entropic, Google, Gemini uses HTTP two plus SSE with their own choice of schema, whereas GitHub and Cursor, Windsoft took a different path. They started using WebSocket, SignalR, and Cursor and Windsoft is using connect with the a product of binary encoding on that. So what is the standardization that we are looking for? Pretty much the transport layer is standardized. What we are looking for is really the last three rows. We want to squeeze it and make it as a schema, which can work on the SSC. So who benefits from this? Right? So entire ecosystem. AI ecosystem benefits from this. We have the middleware integrations like long LangChain, Lite LLM, PortKey. They provide a unified experience for a cuss for the enterprise for to access all the LLM service providers. Now they have to maintain the mapping for with this all these serve service providers, and that's a nightmare to maintain because none of the schemas are easy to integrate with. And now if a new LLM service provider comes in, it becomes a integration nightmare for them to make available themselves to the customers because now they have to integrate with all these LLM vendors, all these middleware vendors. It is also important to note that there is a compliance editing auditing. UAA act mandates that at the customer end, you need to have the auditing done, who generates what, which model generated, and how many tokens were consumed. And everything lays in the payload. So you'll have to look into the payload to figure out and make it available for auditing. Right? So now that also means that if you're using multimodal, then you'll have to have an integration and the mapping with all these 11 vendors. So it is also important to note that, you know, regulatory there could be an enterprise which are regulated, and they can't have an middleware integration with a third party. And they will have to do at their end, and that becomes difficult, and that's exactly what we are trying to highlight. It is also important to note that the HTTP proxies, they can also they will also have to maintain all these, you know, the mappings with whatever the JSON schemas they have. The proposal here is to keep the SSC as it is, standardize the schema with type length and value format. Length is required because currently, HTTP, one HTTP frame can carry multiple payloads or one payload can be carried across multiple, frames. And this is important to note that, you know, we do a byte by byte scanning, to figure out what is the end of the what is the end of the payload. So the opportunity now is for us to standardize it because, you know, the agent to tools and agent to agent to standardize. It is just that the, you know, the server to the client, if it gets standardized, the entire ecosystem gets standardized. What we are looking for is authors who can help us to take it forward. Thank
[00:38:33] Chairperson: you. Alright. Michael p, practical cybersecurity. Oxymoron? We'll find out.
[00:38:43] Michael p: Hi, everyone. I hope I won't take long. This is my only slide. I just wanted to use this opportunity to encourage you to come along to a side meeting. We're organizing on Wednesday afternoon titled practical cybersecurity. So I work for the UK National Cybersecurity Center, and we have a a really good view of the changing cyber threats across across a number of different sectors. We'd like to share some of those some of that insight with you. However, we're really aware that to to defend against the wide range of cyber threats there are, It requires a really broad perspective, and so we'd encourage you to come along, be part of that discussion, and share the insight, and and kind of collaboratively upscale as a as a community. We held a similar side meeting in Montreal, so there's a recording of that and some summaries available on the mailing lists. We also have a practical cybersecurity mailing list if you can't attend. We'd be really keen to to share share insight, so do come along.
[00:39:48] Chairperson: Fantastic. Thank you. James. Oh, you've got a fun last name. Olurandari? Yes. Oh, we've made it challenging to get to the mic this time.
[00:40:16] James Kunle Olurandari: Thank you. Alright. Hi, everyone. My name is James Kunle Olurandari. So I want to present on this, but there is a little trick which I want to mention. So the issue is this. We want to look at the impact of AI, the agentic AIs, right, on resolvers, so to say, so not on DNS. And this is going to be an interesting topic knowing fully well that AI has come to stay. There's no gains seen about that. And the the earlier, the better we start to look at how we can study the way it impacts, you know, the the Internet ecosystem. And I think it's going to be a very good one if we start from the impact of agentic AIs on the resolvers. And, yes, we've not really done so much in that regard, but I believe that with the experts here, we can actually come together to start looking at this. And that is why I'm making this presentation to invite you as experts to see how we can come together to look at the impact of agentic AIs on resolvers. And as a matter of fact, that is going to, you know, lead us into other interesting discoveries by the time we do that measurement. And per a venture, we may later look at this. Maybe, maybe not, but I think it is very good for us to start to look at the impact of agentic AIs on resolvers. And, yes, if you are interested in this study and per event, you have one or two things to add, you can actually contact me through my email address. Oops. I'm not even moving the okay? Anyway, but I think the message has been passed across. So you can actually contact me through the email address. This is just an explanation of how Resolver works, and I'm sure some of us are actually familiar with this. So the main message is this. You can contact me if you're interested in this work, and let's see how we can do the study together so that at the end of the day, we'll have, you know, empirical data that speaks to that. And of course, that is going to be something adding to the body of knowledge within the ITF ecosystem. I think this is something noteworthy, and please please join this course. Thank you very much. That is my email address, and that is my name if you're interested. Please join us.
[00:43:13] Rahul: Thank you.
[00:43:20] Chairperson: Alright. Thank you. Halfway through, Mike Jones, KYAPay:.
[00:43:28] Mike Jones: Hello. Thank you for being here. I'm Mike Jones. I'm going to talk about some interesting work I'm getting to do with one of my clients on agentic commerce in the web we actually have. Not a new web, not new APIs, The web we actually have. So the goal is to enable agents to do things for you on your behalf with your permissions and for the sites that the agents and the APIs that the agents are visiting to know that it's you and to let it through because of that. As opposed to your agents being blocked by the ubiquitous are you human tests that we see all too often in our human interactions with today's web. So, the way that happens is there's a token, which I'll talk about in a sec, which contains a representation of the person authorizing the work and the agent and the parties behind the agent doing the work and possibly where it's allowed to do the work. Standards are a necessity if we're going to have interoperable implementations of these things and I like standards. So at the catalyst buff four months ago, Ecker asked nearly all the presenters the same question. Do you have a bunch of people that I'll I'll do my Ecker speak. Do you have a bunch of people who want to do the thing that you're wanting to do, in which case you need a standard? Or you're just presenting your project, in which case why are you here? Some people weren't very prepared for this question. I'm pleased to report that in the case of this work, there are a bunch of people doing this. It's in production deployment. It's a pre standard but Skyfire wants to make these things a standard with all of you recognizing that you throw something over the wall to the ITF, it's not going to emerge unscathed. It may emerge better. So the core data structure is something called a KYA pay token which represents it's a JWT. I know some things about JWTs. It represents the agent identity and your identity and possibly payment. How do you use it? The primary way you use it is that there's an HTTP header that passes this JWT so that web infrastructure such as CDNs and fraud managers, bot managers, and the whole plethora can look at this common representation and make admit, no admit decisions for your bot. So there's specs. I know things about specs too. The core spec is the token which has been written up for a while. But there's a whole bunch of related specs about how do you use it, how do you exchange it for an OAuth access token. There's some additions to very basic stuff, the AMR claim. It turns out, and this is a fun topic by itself, authentication has evolved a lot in the ten years since we wrote the AMR methods. And we are going to update that hopefully. Why am I here? I'm here to ask for collaboration. Back to me.
[00:47:42] Chairperson: All right. Steven agent accountability composed.
[00:47:53] Steven Meyer: Hello, everyone. My name is Steven Meyer. I'm with Action State Group from San Francisco. I'm alright. We've got some other Californians here. I'm also joined by Tom, Sato, Eman, and Sanbo. I'm just the one doing the talking. We're all about composing for the agent accountability. And the whole thing about what agents are doing, as we know, there's lots and lots of agent talks, is what go when something goes wrong, who you know, how do you know who what happened and who decides that? And so what we do is we have an action digest. And so there's lots and lots of verifiers out there, and those all compose. The can is the authorization. The who is Sambo did the identity bindings, and, Tom did the, receipts there. The what is the agent action capsule, which is, out there as a draft. And audit, of course, is, where SCIT and, rats as well as, the work that Tom did for, tying that together. And so, we have a, we show how the seams all compose. You take your own verifier. And so you have many roots of trust, but they all go into when that decision was made as well as the evidence that's coming from that. And, of course, when that breaks, then you know. And so that gives more trust. And my view is that as you have more capsules all ledgered into Skit, then you can get reputation. And so we invite you to join. We're here about to increase the awareness of of this area that's all emerging. And there's no one working group that's handling this at this moment, but we'd love to work together on the digest. And instead of having lots and lots of bindings that all have to be verified, instead you can just verify to one standard capsule, and then that can be carried across organizations. So that's the whole talk, and really excited for sharing this today. Thank you.
[00:50:05] Chairperson: Thanks. Alright. Our next presenter is remote. H t p HTTP events query, Rahul. And I think we control the slides, so tell us when you need the slides changed.
[00:50:17] Rahul: Do you hear me okay first?
[00:50:20] Chairperson: Yep.
[00:50:21] Rahul: Good. Hi. I'm Rahul. This is HTTP events query. Next slide, please. So if you've been online lately, you've probably been using HTTP, and it's a testament to the architecture of HTTP that it has scaled with the growth of the Internet over the last thirty five years. But there is one set of applications that HTTP does not natively serve. These applications that require real time updates, chat, sports, tracking, monitoring, just to name a few. And because of this, web programmers are forced to implement messaging systems on top. These come in two flavors. You can the first strategy is to switch to another protocol. And apart from adding a burden of implementing a new protocol apart from fetching your data, there the there is increased complexity because all of these protocols walk back one of the good properties that allow HTTP to scale. One blinds the network, one puts a broker back in the middle, one demands clients become server. So push is real, web needs it, and we need it to work at Internet scale. The other option is it's to send notifications over HTTP. These come in a class known as service sent events, which I like to think of as the model t of notifications. You can get it get your notifications over any media type as long as it's the text event stream. And the event source specification itself does not even allow custom headers, and furthermore, it's unmaintained, so good luck extending it. And all of this has resulted in a lot of aftermarket hacks to achieve this functionality. There is clearly a demand for an HTTP based solution. So, therefore, I'm proposing events query, a principal approach to doing real time event notifications over HTTP using the newly standardized query method. Sorry. One more slide. Events query is predicated on the idea that events are the most principled source intuitive source for getting notifications. It discourages the use of endpoints. You can because you're using HTTP stream, you can get reliable and ordered notifications. You can fetch your notifications with the representation in a single request. You can do connect, and you allow intermediaries to participate. Next slide, please. So you discover notifications by looking at the accept query header, and if it's gives you a media type, you know that you can do use it to request notifications. Notifications. Next slide, please. This is a fairly involved example of notifications. Here, you are asking for a media negotiating the media type for the events as well as the media type for the representation. Next slide, please. The response immediately, when you send the request, you get a response back, which will send you the representation with an incremental header which says don't buffer the traffic and a duration. Next slide, please. Now when an event happens, a notification is sent to you. Next slide, please. Another event happens, and this time it's a delete event. So the stream has ended, and the notification, therefore, closed. Next slide, please. So we already have two demos. I would request if any one of you is interested in implementing to please join us. Next slide, please. So you can contact me by connecting on the HTTP API mailing list or opening an issue on side. Thank
[00:54:34] Chairperson: you, Rahul. Alright. Mohammed, confidential confidential computing.
[00:54:42] Mohammed: So are you starting the time? Okay. Yeah. Cool. So welcome, everyone. For the last few ideas in this same hot RFC, I have been proposing the research group, and I'm going to continue on these lines. Basically, confidential computing and digital sovereignty is the current topic that we have, and there has been latest news in the last week if you have been following the technical news. So sovereignty in the cloud is an illusion. Similarly, the handshake that can't keep its promise, why confidential computing's flaw changes the data sovereignty conversation. And confidential computing's trust mechanism is broken. The fix may not exist. All of these are indicating towards the work that we recently did, which we have been actually telling a lot of working groups that, hey, the mechanisms of attestation are actually very much broken. Nobody believed us. Now we have a proof for that. Anyway, we believe that this is not a problem to be solved in the working groups because not many expertise exist in the working groups to actually solve this kind of a research problem. So what we are proposing basically is a research group rather than the working groups. And the potential core problem behind all of these news was that there is a systematic lack of assurance of accountability in confidential computing. The hardware vendors are basically saying, hey. Look. This is the problem of the cloud vendors. The cloud vendors are saying, hey. Look. This this hardware identity has to come from the the chip manufacturers. And then we, in between the researchers, industry, and all that, so so we are in the mess where where the the thing is in the no man's land. And we really want to solve this problem, and this is what the research group is supposed to be. And the sovereignty problem, I have a couple of examples here. For example, the the thing that was sold or advertised as the sovereignty, which is to say that the European sovereignty is broken in a sense that the chips are manufactured somewhere else. It's maybe in The US or somewhere else. It's not like you could have the supply chain attacks and it has already happened. Then The USA, which is supposed to be sovereign, let's say, is also broken in a sense that the connection can be relayed over to anywhere in the world, and we have a CVE now which exists, which I will also describe in the next slides, that there is a complete proof of concept now available that there is a possibility of this evidence going anywhere in the world. Your workload could go anywhere in the world. And from USA perspective, it could be in a country that you are in war with. So this is really this is no sovereignty at all. So who is sovereign in this situation? Actually, the attackers are. And this is the whole idea of the research group that's being proposed. And the IRTF chair was very interested. Last time I proposed it to him and showed him that there was interest in the side meeting, He came up with a very nice question that I really appreciate, and I didn't have the answer to that question in the last. IETF, he asked very nice, people are interested, very good, but why now? Why not, like, let's say after two years or something, or why didn't you propose it like before two years or something like that? If I remember correctly, that was the that was the main argument. I said nice that I don't have an answer, but I do believe that I have an answer now, which is to say that this is the history of all the attacks that have been done and published in the existing literature, except the last one, which we found. 5.9 is the maximum CVSS score standard vulnerabilities, which is the medium. And now what we have is 7.5, which is the high severity. We have three CVs under responsible disclosure, which are actually 9.1, which is to say critical severity. We have two more which are also of 7.5 under responsible disclosure. So a lot of things are actually broken in confidential computing which need to be fixed. So in that regard, what we are looking for are collaborators which will help us scope this research effort for the research group, and we welcome all contributors, and this is the information for the side meeting link to resources. And I really want to thank all the people who have contributed in as authors and contributors. Thanks a lot.
[00:58:55] Mohit: Roland.
[00:58:58] Roland: Yeah. Hi. I'm Roland. I talk about Kira, which is a scalable zero touch routing architecture for resilient control planes. So it addresses an often overlooked problem, namely your control plane connectivity matters a lot. So there are very prominent examples for of outages that happen even at large providers that typically have very good management over the network. For example, you can look up the Facebook outage from October 2021 where they misconfigured BGP routes and that took the whole control plane off. So they were not able to fix their configuration mistake. They had to drive to the data center. They were not able to enter the data center because also the access control was off. Situation was really not so good. There's a paper from Google from 2024 saying basically, yeah, now our infrastructures are so complex, even when we have lots of auditing checks, testing, and what have you, they are interacting with components in unanticipated ways so that outages cannot be actually prevented. They give us statistics for outages of there before WAN, and more than half of the 41 largest outages are actually caused by control plane issues. Surprise, AI won't help you in this situation because once you lost your control connection, you cannot go back to the box and fix it, and even AI cannot get to the box and fix your configuration mistake, maybe. So the objective is to never lose control over your network infrastructure, even if AI is not taking over your network management. Kira strives to provide resilient control plane connectivity. It is scalable, so it is supposed to work with hundreds of thousands of nodes even in a single domain. It is zero touch, and therefore, it doesn't require any configuration, so you cannot break it by misconfiguring anything. It works well in many different topologies. It even works with mobile wireless networks. And we also added some services that are quite useful, like efficient topology discovery even for very large networks within a very short amount of time. You can actually have a distributed key value store that you can make use of in order to let services and components rendezvous so that they can find each other and publish their services. Why am I here? Keyera is basically then most useful if it's embedded in every network device. Would be good if it's an IGF standard. There's an Internet draft out there. So in case you have comments, please send them to the routing area working group list. We have running code, so that's not just some academic work. It's basically running, providing zero touch IPv6 connectivity. I'm looking for maybe operators who would benefit from this solution, collaborations and also implementers. There will be a public site meeting on Thursday, so if you're interested, please come there and maybe we have more time after dinner or doing the dinner together. Thanks.
[01:02:47] Chairperson: Yujia Gao, problem statement and road map for BGP flow spec monitoring.
[01:02:59] Yujia Gao: And hello, everyone. I'm Yujia from lab, and the topic of my presentation is about the problem statement and the road map for BGP flow spike monitoring. And as we know, the BGP flow spike is widely used in camera networks for traffic handling, steering, or some other attack defense. But the the existing flow spec only provides one way control where successful rule propagation does not imply effective enforcement. And especially in data plan, the enforcement outcomes are not directly visible to the controller, and the roles may become may also become ineffective due to policy conflicts, installation failures, or some other troubles. So we think the operator needs execution monitoring to understand the real enforcement behavior and continuously optimize traffic handling strategies. To achieve end to end flow spec monitoring, we identify five key information categories, which has a route life circle, partial realization, enforcement status, traffic treatment outcome, and the correlation information. For this key information, the operator can getting the can analyze the state of the role and getting the faster faster results of the of the device. And here is the whole framework of the flow spec monitoring. It in it has three different different plans, which has the intense plan, control plan, and the data plan. And to achieve the telemetry with different information, we need to doing some extension work to different protocols such as BNP, epiphax, or the young based model. And for this different technique technical components, we search on the existing draft in ITF and funding their half a gap between the existing protocol and the deployment of the flow stack monitoring. In the table, the circle means it's only half a it have a required gap or only only partial support by the existing function. So we so so we think there there have a lot of work to do in this topic And to giving a a same problem space of different drafts, we're trying to write a draft about the road map in the FlossBack monitoring and the proposed date in the IDR working group. And we're also doing some work in the grow of set and other other working group. And in this hot RFC, I'm hoping to find collaborator to refine the road map and the related protocol work and hoping to invite earlier implement implementers. And also hoping to collect feedback from the different working group. And if anyone interested in this topic, please contact me. And I will also have a presentation on Friday in IDR working group. And thank you so much.
[01:06:55] Chairperson: Great. Yuning Zhang. Let's talk about intent.
[01:07:09] Yuning Zhang: Good evening, everyone. I'm Yuning here to present the joint work with my colleagues from Huawei. So Michael actually has kindly shared a lot of use cases where the agent could perform the action on behalf of the user. And one question I want to ask is that would the executor or the execution end point know whether this particular request has been approved by the user or not? So, usually, if we follow the current practices of today, what we could do is to authenticate the agent to validate access token or to check the scope of that token. But those checks do not answer the question or do not necessarily answer question whether those requests were originated from or approved by the user. And one proposal we have today is to have admission point. This is to do the origin check and to check the applicable policy and to check the user constant if applicable before the request becomes net action. And let me explain why the existing token usage or token check are not enough. So if you look at a simple simple example here, so the user could request to refund order 123, but just up to $50. And between the the user initiator and the executor, somewhere between that, there could be a malicious agent that is able to modify the request to refund order one to five and is at least $5,000. So the agent could actually deliver the access token to the executor, and the token could be bound to the agent's key and is everything correct, and the agents could be authenticated fully. But still, the executor refunds $5,000 to the wrong order. And this is because the existing practice do not really answer the question, who approves the request? And this is this distinction is even more important in agentic system because in agentic system, the party who presents the token or presents the request is actually different from the party who originally sends the request. So our proposal is to have the admission point. This is pretty simple between intent originator and the execution endpoint. So intent originator could be a user, an application, or an agent. And in the admission points, four things happen. First, the admission point verifies the request originator. It evaluates the applicable permission policy, and it obtains or verifies user consent if that is applicable and the signs admission decision. So this results in intense animation assertion or IAA is kind of a token. So the agents could carry the original request together with the IAA to the execution endpoint. So this endpoint could further verify the signature and checks the audience and presenter and decides whether it will act or rejects. We have tried to standardize and define the bindings to the IA, but that could be further discussed. We also have implemented a whole flow to explore the belly phase to operators in different use cases. And what we want is collaborators. They want to understand, like, where is the best place for this intent animation work. Thank you.
[01:11:19] Mohit: Great.
[01:11:20] Chairperson: Alright. Shailesh, multi vendor networks.
[01:11:26] Shailesh: Hi. This is Shailesh from Nokia, and I want to introduce a framework for normalizing multi vendor network inputs for LLM assisted network management. So many of us are working on integrating LLMs with network management operations. Right? So one of the challenges that we're gonna face is the multi vendor network elements are going to represent the same data in semantically different languages. Right? And when this data reaches the LLM, because of the nuances in the representation, there might be a risk of misinterpretation by the LLM, which might lead to mismanagement of the network. Right? So just to quote an example, vendor a might say show interface, vendor b can say display interfaces, and vendor c can say interface brief. Right? So all of them are the same intent, but they're different, like, syntax, terminology, field names, and hierarchy, and things like that. Right? So the current systems is that, you know, we have vendor specific data leading to vendor specific prompt design and passing, which is fed into the LLMs, and then we the output of that is considered by vendor specific AI applications. Right? So this limits a lot of things including portability, interoperability, scalability, and prompt reuse. Right? So this is a framework that I propose where we, I mean, I bring in input normalization layer in between the network elements and the central LLM, right, which kind of which starts with the input classifier which says this says that, you know, whether the data is a performance data or a configuration data or an operational response. And based on that, we create a standardized schema, you know, so that is fed into the central LLM. So which gives us, you know, consistent prompts that is fed into the LLMs and a common extraction and, you know, very predictable outputs from the LLM. Right? So I'll be presenting this work up a bit more in detail in the in the margin meeting this week, but I would love to have some feedback from the ITF community to see if this is a valid problem to solve. And I also welcome, you know, volunteers to review the draft and also contributors who can, you know, contribute to the draft, right, and also implement us. Right? Thank you so much.
[01:13:56] Chairperson: Carlos.
[01:14:08] Carlos: Hi, everyone. I'm Carlos from the Federal University of ABC Brazil, and this is joint work with my colleagues, Marcelo and Marco. We we spent the last decade building the computing continuum, also known as the edge cloud continuum. But how do you deploy application logic over it? That's the question.
[01:14:35] Mohit: So
[01:14:38] Carlos: we have, like, we know edge computing, computing infrastructure, cloud native tools like Kubernetes, which is born for the cloud. But how do you do this for the the continuum? We like to to see the the the continuum as the IoT computing continuum composed of different stages from the devices up to the cloud and then down to the user with many intermediate layers. And application logic or services can run on any of these layers. And also, we implemented this in an application for smart irrigation for agriculture and developed and deployed many services. And we had to do this by hand to deploy the services in the specific stages. Also, now we are working on this scenario of drone delivery, and the main problem is collision avoidance. Hence, there must be, like, multiple services spread over different stages for coordinating the collision avoidance process. So there's a bidirectional influence between application development and service deployment. And this is a core idea to to conceive, like, the application as a graph of services, a DAG. And these services, they can be deployed in different stages of the continuum, and they may adapt and and change and migrate like this. And you see this and that. They can run-in different stages. And the idea main idea, the core idea is, like, do you do we map the graph, service graph, onto the stages of the continuum? And there might be, like, different ways, like orchestration, a centralized way, or maybe choreography where the graph is simply sent to all of the the notes and they decide they interact among each other to decide how to to deploy the services or where the services must run. And the question is why now? This is a relevant problem per se, But now with the emergence of AI agents, there are two important things. On the one hand, you have, like, the idea that AI agents can help the deployment process. And on the other hand, the AI agents, they need a place to run themselves. And so this is a draft that is not yet published, a previous version. And I'm I'm looking for collaborators, for reviewers, for implementers. And I'm the proof that it works because last IETF, I presented here in the hot RFC. And now I have new collaborators. Thank you very much.
[01:18:10] Chairperson: Okay. Cheng Wang Gu, fragmented benchmarks.
[01:18:16] Cheng Wang Gu: Hello, everyone. I'm Zicheng Wang from Brain Management Lab. Today, I will introduce our work from fragmented benchmark to interoperable format for evaluating network agents. So for the motivation, we noticed that more and more people, researchers are designing some agents in networking. And more and more projects are building benchmark for those agents. However, there are format test cases, validation method, and evaluation metrics incompatible. So as a result, scenario are difficult to reduce, and the results are difficult to compare. So we propose a network benchmark draft for those agents. In this benchmark framework, it has a common extensible format for network agent evaluation. In this framework, there are four component, the dataset, the environments, and agents, and the the evaluator. And the evaluation workflow has six, seven steps. Along this way, the execution is interactive in the environments, and the results can feed back into the datasets, making the evaluation more accurate accuracy over time. For the task schema, in our bench in our benchmark, one task is just an object with five parts. They intend the topology, initial config, ground truths ground truth solutions, and test the cases. And task can have multiple solutions. Maybe we can set send a different set of commands to achieve the same results. So we want to make a common exchange format for the community. There is another question. Who can we evaluate it? So we also built baseline for this benchmark. We noticed that AI agent may waste many steps on security, timing, and verification. So we built experiment harness. It it has a semantic action interface and a set of reusable skills and make it can make the agent to complete a networking task end to end. I also do some experiment to validation efficiency. The result shows that this baseline can cut off 47% of time and talking use by 81%. It's a recent paper accepted by the APNET conference. So our data draft and the paper is online, and I will welcome collaborators. Thank you.
[01:22:22] Chairperson: Alright. So I think next time we can actually accept maybe up to 25 sessions because we you all are being entirely too efficient and taking away some of my fun in yanking people off stage. But that's fine. That's fine.
[01:22:37] James Kunle Olurandari: I think that's all.
[01:22:39] Chairperson: No. No. With that, I will let you all go. Thank you all very much. Thanks for the presenters. You were brave and solid. Good job.