**Session Date/Time:** 22 Jul 2026 07:00 [00:00:04] **Mike McBride**: Your last time, like, fifteen years ago or something. I don't know when that was. Yeah. It was pretty. [00:00:10] **Stig Venaas**: Yeah. I've been there too. It's nice. [00:00:13] **Mike McBride**: But this time we went to. That's right. It's true. It's a bit of a commitment. Yeah. But it's not too expensive, though. Well, it was, like, [00:00:43] **Stig Venaas**: for my for [00:00:45] **Mike McBride**: the two hour train rip ride, it was just $25 each way, though, each way. [00:00:53] **Stig Venaas**: That's not too bad. [00:00:54] **Mike McBride**: Yeah. It's not too [00:00:55] **Futurewei Representative**: bad. Yeah. [00:00:59] **Mike McBride**: It seemed very cheap, actually. Especially yeah. Is that right? [00:01:08] **Stig Venaas**: Well, certainly, thought the the subway here was expensive, like, €3 something for a long ride on the subway. The Uban? [00:01:21] **Tony Przygienda**: Yeah. Uban? [00:01:22] **Stig Venaas**: Yeah. That's expensive. [00:01:27] **Mike McBride**: That's what I did for here. Bought the. I [00:01:34] **Junie**: bought bought it. [00:01:38] **Mike McBride**: Yeah. For in town, I bought the wine mobile, the end mobile, and then the OBB one for the trip to the mountains. And we gotta start. It's $9.00 1. [00:01:48] **Nils Warnke**: Oh, [00:01:48] **Mike McBride**: okay. So you're up. Yeah. And you can click from here. [00:01:51] **Toerless Eckert**: Yeah. [00:01:52] **Mike McBride**: This is yeah. There we go. [00:01:54] **Stig Venaas**: Click from here. Okay. Welcome to PIM, everyone. Got a nice view this time if the if the presentations get too boring. Let's see. So the note well, hopefully, you've already seen it this week. It's about how we behave in the ITF, so please make sure you are aware of this. This is our agenda. It's actually a pretty busy agenda, lots of topics. Any comments on this? Any changes? Okay. New one. Then I'll go through the working group status. So, basically, all our working group drafts. So the very first one, we are about ready to request publication. We're just waiting for IPR and some boring stuff before we can do that. [00:02:52] **David Lamparter**: So that's part of our work on moving to I g [00:02:55] **Stig Venaas**: m p version three, MLD version two as Internet standard. This is, the host requirements, which is, like, 30 years old and doesn't discuss I p v six or anything. So we're updating that. We got the DR improvement and BDR drafts that are not presented this meeting. Yeah. We'll see how we can make progress on those later. We got two new RFCs. That's our point to multipoint work, so that's nice. Yeah. Let's see. Got some young drafts that we hopefully can wrap up soon, not discussed this time. Lessons learned, we will discuss. Then we got all these, yeah, group discovery mechanisms that are moving on. We got the gap for requesting publication or we it's in last call right now. Then I got two that are in our editors queue, and we got one more that needs to be revised. So yeah. So that work is wrapping up as well. And then let's see. Yeah. The multipath proxy is not discussed today. We got this PFM forwarding enhancements that is in the ISG. There's some discusses that need to be wrapped up, but, yeah, it's moving on. We got the deterministic ECMP. So one slight challenge there is Bill Fenner that brought the work here. He hasn't been available or haven't been able to reach him. I'm trying to figure out what to do if I should just I'm a call for the draft, so maybe I should kinda update it and hope he's okay with the updates. But we need to make progress on that. I think there is good interest in the working group. SRv6 not discussed today, and we are progressing some Lisp attributes to propose standard in in the base draft. That's not presented today, but we will talk about this multicast deployments today, which is kind of why this work. Well, some of us think this works should move to propose standard instead of experimental. And we got the flex algo that was adopted just before last ITF. So that's what's going on in that working group. Any comments, questions? [00:05:46] **Mike McBride**: Did you mention GAP? [00:05:48] **Stig Venaas**: No. It was here. So Oh, okay. [00:05:50] **Mike McBride**: Yeah. Yeah. [00:05:50] **Stig Venaas**: So, basically, it's an ITF last call [00:05:53] **Nokia Representative**: right now. [00:05:53] **Mike McBride**: We're just answering questions. Yeah. [00:05:54] **Stig Venaas**: Yeah. One one slight challenge with GAP maybe is, like, the address allocations, but we'll we'll see how that goes. At least it's with the ISG right now. Yeah. Okay. [00:06:12] **Mike McBride**: Yeah. [00:06:30] **David Lamparter**: Okay. [00:06:38] **Mike McBride**: It's kinda not very convenient place to see the slides. Alright. So we've made a fair amount of progress with this draft, which we've been working on for a couple years now, I think. We thought we are nearing the end, but we received a flurry of comments, and suggestions. So we did add quite a lot. Five new sections towards the end dealing with a lot a lot of different topics. We don't go into a ton of detail. We just add a paragraph or two to kinda share some background and then some basic lessons learned that we've done. The draft has already proven fairly useful just to be able to provide a history to people. So we're not in a big hurry, like I mentioned, but depending upon how many comments we receive between now and the next ITF, we may start thinking about finishing it up. So just to we did add aside from a new section, we did add some new text, which is just basically saying that SAP is typically paired with SDP and added some RFCs per somebody's comments. I can't remember who. The new sections are this is the first of the five. This one first deals with socket APIs and scope IDs, just basically saying that multicast app support has depended on OS socket APIs, not just the protocol itself. And the API design itself is a lesson learned. We put in some RFCs to show how standardizing the socket APIs has has occurred to allow cross platform operation. Mentioned the RFC to deal with scoped addresses, induce introducing the scope ID. Sometimes that scope ID has been inconsistently applied across OSs, and so that's kind of a lesson learned. Textual embedding and higher level libraries have been have helped. The second new section we added is quality service and data plane considerations. Of course, multicast saves aggregate bandwidth, but it does concentrate replication on certain points in the network. And so each fan out router, it turns one incoming packet into multiple outgoing copies. So there is a replication load that we need to be honest about, and that doesn't always scale linearly with the number of receivers. There are hardware replication limits. Consistent QoS treatment across the copies are important and often an overlooked deployment consideration. The third one was control plane scalability and resilience considerations. Large deployments do stress CPU memory and control plane bandwidth. Reconvergence after a link, RP, or topology failure can introduce loss. More RPF checks do depend upon the unicast routing table. And so unicast instability can directly affect multicast forwarding. Just things like that, including graceful restart and just documenting how this can affect multicast. Multicast scoping, we didn't have that in before, and that's kind of an important one in administrative boundaries. So we added. You probably remember how we used TTL based scoping. I p v four introduced administrative scope block, and I p v six did as well. And so we just shared some things. Like, in hindsight, it would have been good to do that from the start, the scoping from the start. Mobility, we're all probably familiar with the bit of the history. I remember we all sharing many drafts, and I think in in here and maybe in the whatever that mobility working group was or maybe still is. But someone suggested we add that. So we added just a little bit about that, some RFCs to talk about mobile I p d six and different approaches to tunneling the traffic through mobility anchors, and it did add a fair amount of complexity. So they'd they saw limited deployment, kind of lost track of mobility, in fact. So I'll get to your questions in a second. So these are we're we're continuing to make slow and steady progress. Thanks to a lot of people, particularly since the last IETf, Prasad and Ahmed provided a lot of draft review comments, so thank you to them. Prasad, you have a question? [00:11:40] **Prasad Miriyala**: More of a comment, Mike. Should we talk a little bit about OEM? I don't know. [00:11:50] **Mike McBride**: Like, a hit the history of OEM with multicast. Yeah. Yeah. Yeah. Yeah. Yeah. Okay. I think we probably should. Stig ahead. Yeah. Let's do that. Thanks. That's a good idea. And Stig mentioned to me that maybe we don't have PGM in there right now either. We probably should talk about pretty good multicast on that area. We talked about that yesterday a little bit in the multicast side meeting. Yeah. David? [00:12:21] **David Lamparter**: Yeah. That was not part of it's more of a note to myself to send you some text. I just realized the behavior of multicast on the loopback interface is also something maybe worth documenting because it does notably differ between operating systems, and it is rather annoying to deal with and probably a lesson learned to not do weird things. [00:12:39] **Mike McBride**: Yeah. Good point. Yep. Thank you. [00:12:43] **Stig Venaas**: Steve here. Couple of comments. Yeah. For the reliable multicast stuff, there's PGM, but there's maybe other mechanisms worth mentioning as well. Do you have pipim port in there? That's another thing. [00:12:58] **Mike McBride**: Don't remember. Yeah. [00:12:59] **Stig Venaas**: So and for mobility, maybe talk about Lisp as well. You mentioned Lisp elsewhere in the document, but that that also solves the mobility [00:13:08] **Mike McBride**: Okay. Issue. That's a good point. Okay. [00:13:12] **Stig Venaas**: More like overlays in general. [00:13:13] **Mike McBride**: Yeah. Yeah. Right. Very good. Thank you. Let's move on. Alright. Yeah. Okay. Are you [00:13:23] **Luis Contreras**: up? Yeah. [00:13:39] **Stig Venaas**: Alright. So this is something I presented last time as well, and I've hopefully made some improvements to the draft since then. Can you hear me okay? I feel like this is a bit low. Okay. So the basic idea, which sounds a little strange perhaps, is that you want to forward native IP for IP for multicast through a an I p v six only core network. So there are people that just for ease of management want to basically do I p v six on the core, but they still do IP for unicast and multicast through that core network. And unicast is handled well by fifty five forty nine and whatever the the fifty five forty nine business is called, but there's no good solution for multicast. And the idea here is you run IPv4 PIM pretty much as usual except since you don't have IPv6 addresses, you will need to use IPv6 headers when you send the PIM messages. And also when you do r p f lookup, you will need to basically find an IPv6 next hop. So so it's kinda like usual, but slightly different. So what's new in this version? One is that it's proposing an all PIM routers IPv6 multicast address for use by IPv4 PIM. So, you know, today, we have a well known v four address, well known v six address, and you're maybe PIM IPv4 PIM process will listen to the v four address and IPv6 PIM process to the v six address. Here, I'm proposing a new v six address that your IPv4 PIM process will listen to, because it wants to talk to other IPv4 PIM processes. And by doing this, we can have total separation from existing implementations, So there's no risk of accidentally having some router getting a strange message it doesn't know how to handle. I think that was one of the concerns last time. Also, means we don't need a hello option to see if our neighbor is capable because a router would only receive these messages if it is capable. Otherwise, it wouldn't listen to this well known address. And so otherwise, the change is all your multicast PIM messages, you know, that are, like, linked local to your neighbors, they will use this new address with IPv6 headers. That's the only change. But it means that your PIM neighbors have IPv6 addresses and, yeah, your your joins, they will contain v four source and group addresses because we are building v four trees. But the upstream neighbor address is a v six address because our neighbors are IPv6 have IPv6 addresses. And to be safe, I think it should be, like, on a given link, you either use this or standard Pim. It's not like you can mix this on a on a link. That could get into a lot of trouble when not all PIM routers seeing each other and so on. What did not change is I still think we should keep RP handling out of scope And couple of things there, for one thing, it's difficult to solve. But, but also, well, RP needs to handle our registers, of course. But if you place the RP outside of your core or give it a v four address, that takes care of that part. If you think about first routers, you can only be a first up router for a v four multicast source if you have an IPv4 address. Else, you you won't be on the same IP subnet as your source. Or the other way around, those IPv6 core routers, they will never need to send the IPv4 PIM register because there's no v four sources. So the only constrict constraint I feel is that you need to give a v four address to your RP, wherever that might be in the network. If you're doing SSM, this is, of course, not a problem. So I'm not sure if I should repeat this from last time, but at least the idea is we have these darker blue routers in the middle that are core routers with v six only. No IPv4 addresses, but they still won't be for multicast to flow through them. Then we have this edge routers that have IPv4 addresses. And if you use RFC 5549 or the successor to that, we will basically have r p four rib that has r p v six next hops, and we'll use that for r p f lookup. And it will tell us how to go through this I p v six core. And, basically, our I p v six next hop is where we send the PIM join, just as usual. It's just that it's a v six address. Yeah. Okay. The last slide talked more about unit cost. This is multicast. I think I'll skip this one. I've I've I've mentioned all all of this already. And and our pin join message looks just like normal except you see the upstream neighbor address is a v six address. And that's easy to do with PIM because we have this encoded Unicast address that has an address family field. And for Unicast, you would need an IPv4 RIB in the core. But if you use RPF vector, we can actually avoid that if you want to. So that means that the only reforest states really in those core routers is the multicast forwarding table. That means, of course, need to forward multicast. So, yeah, hoping for some feedback. Also, yeah, do you think this is worth doing? And and if so, is it ready for adoption? [00:20:18] **David Lamparter**: Yeah. Okay. I guess I I'm asking the same question that you just asked, David and partner. The the the benefit, as you point out, is you can you can get rid of the last I p v four address on the router kind of. To get to have where you have only one, I don't think we need this. And that that is what we need to be asking ourselves here if it's worth doing the work for. I was much more optimistic about this last meeting we talked about this. I I I'm gonna say I'm I'm split on this right now. I I'm not particularly in favor. I'm not particularly against this. It is a whole bunch of fork to get rid of that one last address. And, yeah, that that's that's that's it, basically. [00:21:02] **Stig Venaas**: Yeah. When you talk about that last one last address, you you can kinda do it if you have one IP address, but it needs to use, like, unnumbered with all the same address and all your interfaces and yep. Anyway yep. But, yeah, it's motivated basically by by customers that say, you know, we don't want e four in our network. Any more comments? So but, yeah, talking about code changes, so I have implemented this. Sending packets over v six instead of v four is pretty easy. The the main code change for me that took some effort was the handling of r p v six PIM neighbors and and r p f lookup with v six in prior v four process. [00:21:52] **Mike McBride**: Right. [00:21:54] **Nils Warnke**: As you ask for feedback, we've discussed this yesterday in a in a separate round. And I think the ask of the customers that you mentioned is valid. I also would see some use cases where we are thinking over this solution in parts of our network. [00:22:17] **Stig Venaas**: Sorry. What was the last you said in what network? [00:22:20] **Nils Warnke**: Parts of our network. [00:22:21] **Prasad Miriyala**: Oh, yeah. [00:22:21] **Stig Venaas**: Yep. Yeah. Alright. Mhmm. Yeah. [00:22:23] **Nils Warnke**: Thanks. Sorry for the protocol. I knew it's Deutsche Telekom. Yeah. Thanks. [00:22:36] **Stig Venaas**: Yeah. I can also add that there are lots of networks doing this for unicast today, and and that works. But if you do multicast, you need something like this. Yeah. [00:22:56] **Mike McBride**: So, yeah, we'll, I'll go ahead and ask on the list if, you know, opinions about adoption. It looks like there's a handful of people that think it's ready. Is that fair enough? [00:23:08] **Stig Venaas**: Yep. Sure. Sounds good. Thanks. We'd be happy for more more comments of input later. [00:23:19] **Mike McBride**: Yeah. Issun. [00:23:35] **Yisong Liu**: Hello, everyone. I'm from Mobile. I will present this RPF Vector Conflict Resolution on behalf of my co authors. This draft specified to handle the case where one join message can include the the vector while another that does not so that that that that's a simple scenario. So this draft was firstly presented in the ITF one two three. So the mainly updates from the last ITF to for to now is that the first one is that we had we have added new courses, and And we have added the what scenario for the scenario two, inter area FRR, and for the conflict handling rules for the scenario two. And then we also address the the comments from state. Thanks to Stig. And update the we we clarify that we update the RFC fifty four ninety six and the RFC seventy eight ninety one. By the sublet supplementing the existing specifications with handling rules for these specific conflict scenarios. So next, I will give some introduction for the two scenarios. Firstly is the partial r p f vector. And in this figure, we can see that for the r one for the receiver r one, the primary join, like the the red line in the feed. And from r four, r three, r two to r one. And the backup join, like the blue blue one, blue blue blue line in the in the. And so in r six in r six, for the blue line, we can create the entry incoming interface for the f one and for the outgoing interface for the f two. But in but for the r two, it only had the primary join. And like like the green green line in this week, for the the same the same r six for the incoming interface is if two and outgoing interface is is three. So that's that's that's a complication in this scenario. The Scenario two is for the inter-area multicast FR. So for the inter-area multicast, should carry the RPF vector. So for the for the r two, we can only have the primary join, and we can only contain one vector for the ABR. That's in the r six, we have the entry two that in that the green line in in this big. That's in the r six, we have the incoming interface is if two, and outgoing interface is if three. And for the r one r one, it have had the backup join. The the backup join, like the blue line in this state, and it it will take the two vector, r five and ABR. So it it will in the r six, it's the incoming interface is e one and outgoing interface is two. So there's a a a conviction in this scenario. So summary summarize the that the problem statement is that, when the RFC 5384, when you draw messages with the same attribute types, so priority hits the neigh neighbor address selection. And if the address are the same, selected by the interface index and for the RFC seventy eight ninety one, and then while PIM join message with join attributes is received, the router must verify the consistent order of the RPF vector types. If the order is consistent, the types are order considered equal. Selection is made according to the RFC fifty three eighty four. So neither of of of this RFCs specify how to handle the case we have mentioned in previous slides. So here, we give some recommended handling mechanisms. So when the two conflicting join messages are received, firstly, if one with the RPF vector and the other without the the join message without the a vector is given the priority. The second one is that if one join message only contains a type zero of the vector and the other contains a type four, type four means the explicit object vector. So the join message will this is only the the type zero vector is given the priority. The third one is that if one join only contains a single type zero vector and the the other contains a stack of the multiple type zero vector, So the join message with a fewer RPF vectors give the priority. So in the previous scenario, we can get these results for for the scenario one. The recommended handling process in this that for the r three, we can get the incoming interface is the Ef-one, and outgoing interface is the Ef-two and Ef-three. So the the backup from the r one will be blocked. And the so so in the R six, we have the incoming interface is Ef-two and the outgoing interface is Ef-three. And for the scenario two, the the backup join from the r one was blocked in the R six. So for the r two, we have the the entry. Incoming interface is F1, and the outgoing interface is F2 and F3. And for the r six, we have we have the incoming interface with 2 and the outgoing interface with 3. Only the the the the the primary drawn from r two existing. So that's all. We have come more questions and comments, and our requesters think that the draft is ready for the organization. Thank you. [00:31:35] **Stig Venaas**: Yes. Steve here. One question. So I'm wondering if this is backwards compatible with RFC fifty three eighty four, I think, the original drone attribute RFC. So one thing it says there is that if if two routers send the same attributes, the one with the highest IP address, I think, is the winner. But here, you are saying that the one with the shortest list is the winner. So I think maybe old implementations, they may look at just the first attributes and see which one has the highest IP address. But now they should look at the length of the list. [00:32:21] **Yisong Liu**: Yeah. But because the the the fewer the the fewer IP vectors, I mean, that yeah. We want to we want to assure that the the basic PIM join function would be assured. Yeah. So so that's the we we make the the fewer is the [00:32:52] **Stig Venaas**: Yeah. Yeah. No. Sorry. Yeah. Let me just add. So I I agree. It's seems like the the best behavior. My only slight concern is, if some router implements the all RFC and another router implements this improvement, they might maybe make different choices. [00:33:12] **Prasad Miriyala**: Cisco Systems, so do you think making that this RFC update, the previous one, make more sense? [00:33:20] **Stig Venaas**: Yeah. I think this maybe should update that. But you might maybe also need on a hello option or something to indicate that you support these new rules. I'm not slightly worried otherwise that even if you say the RFC updates, it will be implementations that only use the old rules. Sure. [00:33:53] **Mike McBride**: Okay. Thank you. There's no more questions. We we'll ask this question on the list. You. We did have a poll. There was one person in Okay. That was ready. So other people otherwise, people don't have an opinion. Thank you, Usang. Yeah. Luis, you're up. Yeah. Let me find it. [00:34:26] **Luis Contreras**: Yeah. Hello, everyone. This is Luis from Telefonica. We present this new version of the the this proposition for having a a way of scenery in eco mode using IGMP/MLD extensions. We're doing behalf of Marcello and Benjamin that recently joined the the draft. So just summarizing a little bit the rationale here, the the idea is to enable a a capability for the end users, for the listeners to signal the willingness of having a more efficient or energy efficient delivery of the of the content. But this basically could apply, of course, to the the resolution, for instance, of the of the video streaming that is being received. But also from this signaling, other actions could be derived, maybe optimizing the energy consuming the path or maybe from the devices even to take decisions, maybe going switch off or or whatever. The draft does not enter on on what would be the the the the actions being triggered by by the the signaling, but simply proposes the the way of, yeah, signaling these capabilities so that later on, some actions could be triggered as set for optimizing the energy consumption in the overall delivery of the of the streaming. Yep. So for that but we are leveraging sorry. I forgot to mention on the extensions that allow that were defined in in I don't remember now the number. I think it's it's later in the slide. But these TLBs that can be added to the IGMP/MLD messages for for extensions of of it of them. So from zero zero to zero one, apart from from Harry, Marisol, and and Benjamin in as co authors, we have improved the discussion of the product statements or going further into the details of what would be the advantages of having this kind of signaling. We have other relationship with the energy efficient network management capabilities in both terms. So, you know, where there there is this other working group, Green, that is working on this management of energy efficient management of of the network. So, basically, we provide some hints there. We have also added some discussion about use cases and going very quick through them. Large scale energy aware multicast distribution where, basically, this this echo mode preferences could help to select the delivery alternatives. So, basically, the proper resolution of the video streaming. Energy aware multicast video streaming, so basically selecting the the more efficient path. Mobile multicast environments where maybe the battery of the of the terminals could be benefited as well of enabling this capability of signaling more efficient delivery mode. Carbon aware multicast delivery, this is a a little going a little bit a step beyond the energy efficiency. So taking into account the energy mix that is feeding the power of the central offices and as a probably a little bit more futuristic. And somehow also somehow also exploring the possibility of having ways of incentivizing users for doing more responsible from the point of view of energy consumption. Just in in the media operations working group, also, there there were some other interesting things that maybe could be taken into account. For instance, the distance from the from the people watching the TV size of the screen regarding the resolution. So, also, somehow, this cow this kind of scenarios could be taken into account for enabling the signaling of preferences in terms of energy consumption. So we also in this new version, we have refined the proposition of the TLV extension. I will go quickly through that in the next slides. And we have proposed some some metrics. I will also cover cover them very quick. Regarding the the TLV. Yeah. The the RFC that we are leveraging on is the ninety two seventy nine. This RFC defines this the possibility of extending the messages for including additional information. So we'll, as I leveraging on that for defining this TLV. So this echo mode is intended to be supported in both query and and report messages. So from the listener to the to the network, from the network to the listener. And, essentially, what we are proposing in this new version, we have refined a little bit the the TLV, the regional TLV. We are adding two two fields, the echo level and the preference. I will cover in the next slides. So for the echo mode for the echo level, we are initially define identifying four modes. The no echo mode, that will be basically the conventional delivery of the video streaming, so no action, basically. So it would be so how are it will guarantee some so how they get the backward compatibility with whatever we have now. Light echo mode where basically the energy saving actions are acceptable if service impact is is negligible. So, basically, guaranteeing that the quality of the video streaming is not impacted. Moderated echo mode where basically the the listener asserts on optimization even there could be some yeah. Basically, playing with alternatives for delivering the the the video, and then a a more aggressive echo mode where even the quality could be impacted so the listener express the the the fact that the is okay with the even if there is some impact on the quality of the video streaming being received. Regarding the preference field, we have proposing for values, let's say, the targeting the energy efficiency, targeting the carbon awareness, targeting the renewable energy preference, so basically leveraging on renewable sources. And so another open thing that would be the operator defined criteria that, yeah, basically, will be on the hands of the operator. You see the caution symbol in the slide because these these values are basically initial ideas, of course, open for being discussed and and any suggestion is jump but any suggestion involved on what, basically. Then regarding metrics well, there is a question. I don't know if I don't if okay. So we also propose some some metrics. Again, the caution symbol is just to to to signal that that these are initial propositions, so open to any any suggestion, any discussion, any feedback. So we are considering the yeah. Having some metrics for the counting the receivers that are in this mode. Also, having an understanding what is the distribution of the listeners in the in the network requesting this. Also, for accounting the energy or the carbon that has been avoided due to the action of the customers requesting this common way of the the streaming. And, yeah, also the the the result of the policy has been accepted, ignored, downgraded, etcetera. So going now to the final slide. So, essentially, summarizing with this mode, we we basically trying to enable a way of for the customer to signal the willingness of having a more efficient delivery of the of the content. This could trigger a set several actions, not only in the network, but also in the devices, in the set of box, in the TV, so. But we are not entering in that space. It's just simply defining defining the enabler for that. And just to comment that the document was also presented in green this week, one of our ideas is to trying to to set the link with the green management capabilities so that this kind of signal could be leveraged for for those for the green framework at the end for maybe yeah. Even collecting metrics, even taking decisions and so on so far. And, yeah, we will keep working on that, preparing a new version for next ITF, and and, basically, for today is to to ask if this makes sense for the working group to to keep working on it. Thank you. [00:42:18] **Prasad Miriyala**: Cisco Systems. So the way I'm understanding, the receiver wants to join some eco mode. But if you look at multicast, so you have host who are receivers, then you have source. So source is going to send different content for different eco mode. How are we translating this all the way to whether overlay join or PIM join, any of those l three? Because you have to connect it end to end. [00:42:42] **Tony Przygienda**: Yes. Will it [00:42:43] **Prasad Miriyala**: not make more sense to have different eco mode being a different group itself? [00:42:50] **Luis Contreras**: Could make sense, probably. So some some point here will be that that maybe depending on on the volume of of customers requesting one this eco mode for one group or whatever to somehow adapt to the multicastry to that yeah. To the number of of customers requesting this echo mode. I guess that the also, we need to take into account the the resolution of the terminal and so so probably there will be disassociation in in in the terms that as long as we are we are collecting the willingness of the different users, we can even, let's say, not extend the multicast tree up to the very end if there are no customers requesting that resolution or within with the without the willingness of of requesting that resolution at that particular point. So, basically, adapting the tree to the willingness of the different users spreading the network to the to the resolution that they want. That could be an impact a result. [00:43:44] **Prasad Miriyala**: Yeah. So the reason reason for this question with respect to adoption, so I have seen so this requires a change in the host application itself, which is, let's say, for IPTV set top boxes. So change doing any migration or upgrade to set top boxes is relatively harder compared to coming up with new groups. [00:44:02] **Tony Przygienda**: Mhmm. [00:44:02] **Prasad Miriyala**: So if you just push that instead of these 100 groups now, you can pick out of 1,000. That is much easier compared to you say that you have to upgrade the iGMP protocol stack. So my comment is related to the adoption of this whole work in the real field. [00:44:19] **Luis Contreras**: Okay. Thanks for for the comment. I think on that. [00:44:24] **Stig Venaas**: Yeah. One one more comment. So, yeah, if if you need to basically have different content, like different rate or something based on the mode, then, yeah, the source needs to know about it. So I kinda agree with on that. But one thing is, I guess, other other ways you can, you know, have eco mode is by using certain links in a network, like not necessarily doing the shortest path or something. Right? And then it's something the source is not really aware of. So then the routers need to know whether it's eco mode or not. You could possibly configure statically at the last hop that, okay, this router knows that this this group is eco mode. But it yeah. It's kinda nice to make it dynamic. But I agree it's a challenge upgrading set of boxes and so on. [00:45:19] **Luis Contreras**: Okay. We will think on that maybe, yeah, to see what how we can leverage on this for doing something that will be less impactful [00:45:28] **Stig Venaas**: for the for the. Yeah. By the way, we there there is some very old expired draft that talks about using some iGMP MLD signaling towards the source itself so the source can know if there are receivers in the network or not. If you had something like that, you could actually signal from the first sub router up to the source also and say, EcoMode is requested. [00:45:54] **Tony Przygienda**: Okay. [00:45:55] **Stig Venaas**: So that could allow the source to do some changes based on this mode. [00:45:58] **Luis Contreras**: Thanks for the for the point there. I will I will look for that. Thank you. Yeah. Thank you. [00:46:02] **Mike McBride**: Are you leaning towards requesting adoption here or in green? [00:46:07] **Luis Contreras**: The the the adoption games, I think, will be here. Not in green. Green basically can use this for for figuring other purposes, but I I think this will be the protocol wise to to be handled here. [00:46:21] **Mike McBride**: Are you ready for that adoption call to be made or not? [00:46:25] **Luis Contreras**: I'm not sure if she's so soon. Maybe we need to address the comments that [00:46:29] **Mike McBride**: Okay. [00:46:29] **Luis Contreras**: That Steve and and, like, I might mention and probably come with a more solid thing from the next time. Thank [00:46:35] **Mike McBride**: you. Perfect. [00:46:36] **Luis Contreras**: Thank you. [00:46:44] **Mike McBride**: Yeah. David. [00:46:58] **David Lamparter**: Hello. This is the second time this is being presented here. First time was on Monday in 6man. This is for a new scope for IPv6 multicast addresses to behave a little bit differently. Right. I guess this is not super useful here to talk about updates. I'm bringing this here to make all of you aware to get some input from the the Pim working group on this. And there's also, very notably, completely left open right now in the document how it should behave for MLD. And, yeah, let's just go into it. You're probably all aware of these special purpose Ethernet multicast addresses that have behavior specified by 802.1q, especially the first one here on the slide, the one eighty c two whatever zero e address is the LDP address, and that one is never forwarded by an Ethernet switch. You can find all of these in the 802.1Q spec. These three ones here on the slide are special in that they don't have special in in that they don't have protocol behavior attached from IEEE eight zero two. And I'm looking at this because there's been multiple drafts in other working groups at the IETf that were looking at using LLDP as a discovery protocol. And the main benefit that you get with LLDP over any other discovery protocol is that you can rely on the fact that it never goes across the switch. And I would like to point out this also means that it never goes into a VXLAN overlay or something like that. So if you want to do some signaling between a host on and a top of rack switch that is putting everything into VXLAN that you're also doing with the host, then using these groups is also useful to do something that does not go into your whole overlay network. And rather than having to go to LDP to do all of this, especially if it's routing information, and I'm gonna note the middle of these drafts that does just plain routing over LLDP, I rather do have an IPv6 option to do that than be forced into mess to integrate the layers like that. The TLDR version of what this draft does is these three multi actually, multicast addresses get mapped to these three Ethernet addresses. The draft is specifies a generic mapping. It's it's not these three specific addresses, but these three addresses are the ones that make sense and that are relevant here. And that's the thing really I'm trying to get to. It's an update to RFC 2464, which is the original specification that says how you put I p v six multicast or ethernet. You probably all know the thirty three thirty three prefix. And this this is an extension essentially to the rules given there, which is that the very simple rule that you just use the the last four bytes and put the 33 in front. And this says, okay. And for these special groups, you do this different thing. Yeah. It's a new range at the top level of the IPv6 multicast address registry. The bit combination is currently undefined. It just makes no sense. So this is using a new top level multicast space, I guess. It does put a registry for the layer two types to use with this because even though right now the document only speaks about Ethernet, at some point, there might be another link link layer technology that has completely different groups specified that have completely different behavior, and those will have different addresses in this out of necessity. So it's it's living room. There's only Ethernet right now. I I mean, we all know that Ethernet is 99% at this point, so that's the main thing. And then to go with that, it just specifies the mapping rules for these addresses so you get the special purpose Ethernet multicast groups. Yeah. This is the table from 802Dot 1 q with the three groups that I'm highlighting here. I do want to point out some details. This can only ever be less than the entire I v v six link because if it were somehow more, that would be the link for v six. It it just makes no sense by by definition, I guess. I want to be very clear that this doesn't change behavior of these groups in any way. This is only exposing existing behavior, and I'm noting these three groups specifically because they are not specified for use with exactly one eight zero two protocol. They are defined as general behavior, of, Ethernet bridges switches. And this the the the document tries to be, reasonable in not, trading on, I on the IEEE's scope here, and say that you should think about what you're doing and maybe not use some other group that eight zero two is using for something specific. This does also kind of define how this relates to to the IEEE. This is I p v six making use of an existing thing. It will get a liaison statement if it goes forward, and we'll we'll hopefully get some input there if we need it. It is relatively simple as an idea. And yeah. The main open question here is, what hosts should do for MLD on this. It just says nothing about this right now. The document is just left as a to do. As I've already said, the forwarding behavior for these groups is completely out like, changing anything there is completely out of the question, and that must be absolutely clear. This exposes what would happen with these Ethernet multicast groups by default and nothing should affect that. I'm a little bit concerned that on some devices, if they implement MLD snooping, they might additionally restrict behavior for these groups only for v six packets and or I hope I really hope they don't. They might, in the very worst case, affect behavior for these groups. The only thing that can be done for that is just to put in a draft that you're not supposed to do that. The my best my best, well, estimate of what what this thing to do here is what I've put at the bottom here. Hosts may send reports. They're not required to. They are not required not to. There's no point in querying. This is not going to be routable in any case. And since snooping makes no sense, the whole query is essentially pointless. And it I yeah. I I do think it should have the must not on on doing snooping on these groups and especially on changing behavior for them. That's that's the main thing I'm I'm hoping for input for at this point. Yeah. Lastly, yeah, the idea is truly trivial. I implemented this during the hackathon in a few hours. It's 31 lines of code to get into the Linux kernel. If you're interested in more details about why the I t f one twenty four 6man presentation has a bit more on that, and I think that's it. Any questions? Not everyone at once. [00:54:58] **Mike McBride**: Well, thanks for presenting here. Thank you. Thanks for your time. Yeah. Right. Have a schedule. We're doing fine. Alright, Prasad. Is that right? [00:55:20] **Yisong Liu**: Yeah. [00:55:32] **Junie**: Hello? [00:55:33] **Mike McBride**: Yeah. Am I audible? We can hear you. Yep. [00:55:36] **Junie**: Okay. So good morning. Hello, everyone. Yeah. Today, I will be presenting the observations tied to the draft list for multicast deployment. Yeah. Rather than working through the document, I'd like to share from the operation perspective as of what we have actually seen running IPv4 LISP multicast in the production during the large scale network infrastructure out in our environment. So yeah. I and I suspect that operator input here is relatively rare in this working group. So just consider this as field report, and I will stay at the protocol protocol level throughout. So next slide, please. Okay. So, yeah, the deployment context involves multiple list sites and grouped into the campus. So, like, campus a and b here on the diagram. So they are connected through an IP underlay with a list overlay on top. And two things worth mentioning. First, we run both layer two BUM traffic forwarding and the layer three routed multicast side by side. So we suppose ASM, so any source multicast and source specific multicast in active views. And second, the underlay itself does only native multicast, so we are not doing any head end replication anywhere. So it means the combination of a native multicast underlay carrying out both layer three and layer two overlay traffic. So it's the background, for everything. So, I will talk about. So next slide, please. Yeah. Let let me start with some good news. Yeah. So fur first, so during the receiver's PIN join event, So encapsulating the PIN inside a Unicast list packet towards the upstream hops, so proven to be a simple and reliable mechanism. And second, the pin drawing attributes also give the receiver egress to the router an effective way to signal the underlay transport information. So back to the ingress tunnel router. So it's on how how the overlay is signaling or works here. And the third one, a choice has been made to collapse the underlay down to the source specific source specific only multicast only for overlay layer three multicast. So this also reduced the control plane complexity. The force would be the RP RP placement. So we we have a stock with static RP and any cost RP and MSDP. So not flashy, but a stable way or no combination, so for for years. And finally, this would be a more of a operational point than a protocol point. So pushing underlay and overlay configuration purely through an orchestration tools has prevented from human errors. That's a very important for the operation. And next slide, please. So here, I would like to zoom in on two specific design choices. The first one is what we call n equal to one underlay group mapping for boom traffic. Concretely, it means many layer two list layer two instances. So let's say m f m of them. So all mapped done to a single underlay multicast group. So this is simple to run, but trade off is less traffic separation. So all these so it means all these layer two instances share one single underlay group. And every egress tunnel routers, you see boom traffic set, so probably not necessary target for themselves and has to filter it locally. And then the second one is the overlay to underlay group mapping across the sites. This is strictly required to prevent multicasted loops when traffic across list sites. And and this operation requirements that come with come with it is that the mapping has to stay consistent across the every single devices in the list domain. So when, yeah, when one of it get a run configured, and then you can end up either a loop or even silent traffic loss. So next slide, please. So, yeah, here brings me to where things get harder. So the first one is troubleshooting. Because because the state spends now two distinct layers, so overlay and multicast overlays. So considering that many operators already have much less experience with multicast routing than unicast routing. Yeah. And this dramatically raised the bar for the network engineers. So so in practice, only a very small number of people in the team can actually, yeah, reason about both layer at once and color rate what they have seen. And it's in some scenarios, it's very difficult to have a full picture of what went wrong. So and this also loops back to the previous slide that because both layer needs to be configured together, the deployment ends up heavily dependent on automation. And second one is the specific concrete issue we encountered in this push to talk style scenario where every new call initiated locates a new multicast group. We also observed that that the join the join simply hasn't converged yet by the time the source has started sending, which leads to the receiver missing the call entirely. So this is, yeah, the real gap I'd like to I'd I'd like to working groups input on. So maybe is there any recommended protocol pattern or perhaps a faster drawing signaling or any other maxing mechanism to avoid, this joint latency loss for short lived sessions. It it could be, enhancement or improvement, to the protocol. So yeah. So to wrap up, yeah, at least multicast holds well in production, but this two layer troubleshooting and the short lived session converges remain very challenge. I I I also suspect we are not only not only the not not the only operators who have hit these ages. And yeah. And that's that's all from my side. Thank you. So welcome for any questions. [01:04:18] **Mike McBride**: Thank you, Junie. So this is something that you, I'm assuming, also presented in Lisp, and you're probably gonna seek adoption in Lisp. Is that correct? Yeah. Okay. Any questions for him? No. Okay. Thank you. Thank you for your time. Yep. No more questions. Alright. So this and the next draft are spillover from a multicast for AI side meeting that we had yesterday, and we've we've had three of them now. If you're not familiar with this this side meeting, it's it should be something that you should look into it. The my meeting was recorded, and it's kind of a big deal because there's a decent chance, in my opinion, that it will become a working group, and it has some pretty wide ramifications for AI in general. So and the PIM working group should be agnostic to what's the solutions are and the different proposals. So that's why we decided to present them here. If it ever does become a working group, then it it would be likely being dealt with there. But we would need to have this group and beer and other working groups opinions as well. So just kind of the kind of a quick overview if this is new to you. So the predominant traffic pattern within AI is point to multipoint. So whether that's in training or at inference when you're asking actually a question to a model on cloud or whatever, the traffic pattern is point to multipoint. And that's being handled today by Unicast. And so the reason that we have had these side meetings is that we feel it would be much more efficient to use multicast for these point to multipoint traffic patterns, especially as the scale increases to, you know, starting to use, you know, tens to hundreds of thousands of GPUs that are having this traffic pattern. And so, you know, just to help you understand how it works, like, at at inference when you're asking a question. Let's say you asked a question that was, like, 20 words, and that would be converted into, like it's like it's like two thirds of a token. So I think it's like you know, anyway, so let's say that that's converted into a certain amount of tokens, whether that's 30 tokens or whatever. Each of those tokens for each user is sent to a different expert if you're using multi of experts. And if you have millions of users asking these questions, then that's a lot of point to multipoint traffic. And so that's that's the purpose of why this side meeting is happening. So these questions were have been asked during the side meeting is, you know, what do we do? Do we use an existing solution, whether it's PIM or beer or create a new protocol that's specific to AI to be able to handle that. And so these are kind of the questions, and the purpose of this particular draft is to evaluate these different solutions. Even some other ones came up yesterday. It's like, maybe we can use some new lightweight method. Maybe we can use PGM to do something like this. There's there's a lot of different options. In some scenarios with n I AI, you can that are stable, that are going to all recipients. You know, like, if there's a situation where all the GPUs have to respond with their correction slips, so they're called gradients, they need to be sent somewhere, then you could use PIM to do that because the it's just it's it's not dynamically changing for every token. Or if you're distributing in a model to certain locations and those locations are static, then you can use something like PIM to do that. But for the most part, in dynamic scenarios, you gotta use something else. And beer is the leading candidate because you just there's no state. You just put in the recipients into the packet, it goes to where it needs to go. So I've kinda discussed this just a little hand wavy here, but this is that's kinda sets the stage of why we are discussing that. And it's a in my opinion, it's fascinating. It's something that we'll probably be working on for some time. And it does take a lot of it will eventually probably take a lot of coordination with other standards bodies, I'm gonna guess. It was brought up that the Ultra Ethernet consortium is also working on these kind of things. You know, we do our thing here for, you know, protocols, layer three protocols primarily. But there will be need need to be some coordination with other standard bodies as well if we do wind up doing this. I think I mentioned training collectives. They, as I mentioned, do have stable, membership. So there are you know, there's different things that we can do with existing protocols and maybe, creating some new ones as well. Let's see if there's anything that's missed here. Yeah. We'll just keep going here. So let me just make one comment here. So InfiniBand is the is what's used today for the solution to have AI work. And there is this baseline transport header, b b t h. And so if we need to have this, you know, new solutions be compatible with InfiniBand, then we would need to likely include the b t h in our solution. So today, what we could do is put that b t h in a beer header and just forward that as well. So that would be a good candidate for that. And let's see. Congestion control is an important aspect of this Scalability, as I mentioned, as that continues to grow, that's why we you know, there's there's a lot of need for trying to make this more efficient. So this is kind of where the the interesting part is is and, again, what I alluded to at the beginning is that we extended a existing protocol like beer for this mixture of experts where you have a which is the predominant traffic pattern where you've instead of using the entire model to answer a question or to do the training, you just select, you know, maybe 10 experts to handle it. But, again, those experts are chosen per token. So it's just a lot of this from one GPU to these 10 other GPUs, and then it changes for every other one. So any question you ever ask, it's typically always using this MOE pattern. So we do extend beer to Bell support m MOE. And there are drafts that have been presented in BIER and elsewhere to support that. So there's arguments for doing that. It's an existing standard. It's multi vendor implementation, faster path to deployment. It doesn't you know, changing which receivers get a packet doesn't require any control plane signaling. Unlike the tree bay tree based mechanisms, it's well suited to dynamic membership, like MOE, like I just described, and the header is payload agnostic. And it can carry that BTH, as I mentioned before, without modifications. One of the big points with this is a drawback to all the different multicast protocols is that there's no native ACK aggregation. So you have to have an acknowledgment in AI traffic patterns like I mentioned before. We need to see the gradients or the con the con the correction slips to and then then you send out the correction to all the different participants in the model, and then they update their their weights. And so that's something that's important. And so there have been some proposals to be able to add ACR aggregation to to beer, particularly. The bit string size limits, there is a limitation there. There is complexity when you start using hierarchical domains. There's no defined congestion control interaction. And so these are the kind of things that we're trying to weigh in whether we try to extend an existing protocol or do we create something new. So, you know, there's arguments for creating a new protocol. And, again, this is very early stage. And if if this multicast for AI does become a working group, we probably need to be able to have some clear deliverables of justifying having a new protocol. Because if there's something that if we decide to just extend an existing protocol like beer, then that would just be handled in the beer working group. There wouldn't be a need, I don't think, to have a a a new working group. So you could have a purpose built new protocol for AC aggregation for this BTH RDMA over Converge Ethernet or BTH compatibility that can congestion control. AI data data center is a distinct unique group, and it could just this new protocol could be targeted for that use case. And it could enable co design interworking points. So, you know, do you do the aggregation at the source tour or some other places. You could define a protocol to be able to handle those certain certain areas. But there are limitations. Of course, whenever you have a new protocol, it does require new implementations across ASICs and all that. It needs a new working group charter operators. It would take a while for them to ever deploy it and risk potential vendor fragmentation. So we try to do a comparative analysis in this draft. It's early stages, and we're continuing to develop it. But ACK aggregation is important. Interactivity is poor for all all the existing multicast protocols. Interactivity, meaning, you know, the acknowledgments coming back, particularly, because we have one way trees. There's not a return path even though there are some proposals. Bidirectional PAM, which a lot of us have not even discussed for many years. That's something that's come up as a bidirectional tree topology, but there's no reduction capabilities that AI needs to reduce all these acknowledgments to come in and then just and then forward it on. Highly dynamic group membership requires fast, lightweight group formation, not long leave low not long live tree setup. And so adding stateful per group ACG aggregation to beer could abandon that stateless design. These are some of the things that we have in there. There's some open issues like ACG aggregation, BTH classification, interaction with packet spraying, and and things like that. So and I think we have time. But so this is it'll be interesting to get our opinions here in this working group as to, you know, where should we go with this? And it's, you know, it's a little bit of a mute point because, again, we have these side meetings, and that's where the bulk of this is gonna happen. But we just need to get your feedback, get your ideas of where we should go. There's no existing working group handling this, but it could very well become a working group. We're well aware to handle certain things, certain proposals. We are chartered for multicast protocol development. We could, if needed, the upcoming year, add AI language to our charter if we need to. We definitely need to scope it because, you know, anything beer related, we need to go to the beer working group. Anything routing control plane would need to go to other working groups as well. So the floor is open. [01:17:06] **Toerless Eckert**: Yeah. Thales Eckert. So the egg aggregation that that's kind of the biggest unknown for us. I was trying in the the site meeting already explaining that there has been a whole empire around that reliable multicast transport twenty years ago now. The conclusion there was actually that you're doing neck implosion. That's what they called it. So you sent necks, and you then reduce them. Right? So they actually call it implosion. But I don't I would kind of just to keep life simple for us, go with the assumption that this is not needed and the application people first have to prove that it's needed because it's typically needed when you have a large number of recipients, not if you have 10 or so. Right? So if we're talking about millions of small groups, I would be very surprised if we actually need it. Right? So that's that's one of the things we really need to get out of the people on the application side to see the patterns, what to do. If, you know, you just send 10 next back, that that's that's not an issue, I think. Right? So, hopefully, we can we can avoid having to do that. But there's definitely a a question, what would be the best reliable multicast transport to put kind of in terms of the retransmission and so on. Right? There there is no such protocol for multicast right now, and I don't think that the unicast protocols like what InfiniBand and others are doing are well suited. It's true that on the ultra Ethernet consortium, they're probably having to solve the the the the same problem. So we are going to have solution wise to look into the reliable transport aspect, but hopefully not on the routers. Right? So I think that's my my hope. Right? Because we've been around the block. We've implemented that NAC implosion, and we know how difficult that is. Going back to the network layer. So first of all, the the stateless, more scalable than kind of, you know, what I call now flat bit string as beer does it. We have presented the first draft with, you know, a prototype implementation in Huawei in 2022 in the beer working group. Since then, we've done different variations of that, also built for a variation at p four Tofino implementation together with University of Tubingen. There are existing implementations and approaches to do this, so it would be very good from the proponents of of these application spaces to take a look in there and say what they don't like about them, what what may be needed to be improved. Right? So we didn't progress with the standardization request for that because, again, we wanted to see more interest in this. But, again, this is not a technology free zone that we're entering. Right? So there there is existing work on that. With respect to which working group it fits in and so on or whether it's a new protocol, let's say this in segment routing for I p v six, a a simple list was replaced with a very complex structures, and they just called it compression. And they kept it in the same header that that they'd create in the first place because they liked the header. Maybe in our case, it's the opposite. Right? It's not about arguing about the kind of structure as opposed to a flat bit string, but that the beer working group had done a very limited, you know, view of what encapsulations for the header are feasible for all the different networks with just beer in six. Right? So I think that's one of the things people with an I p v six network have been struggling with, and I think that would be lovely to to be more amenable to to different network requirements. But when it comes to, you know, such a stable structure as opposed to flat bit string, I think the the the first thing really is, again, ask we need to have, from the side effort, a better write up of the exact scale of the the use cases. Right? For example, they have a very nice proposal. I'm not sure if they were adopted in beer of just using different subdomains to scale up to the maximum size that we can already do with beer. Right? So we should be able to show, yep, you can go this way in this use case with beer, but we need to go beyond it. Or, no, we don't. Right? So I I I think that's that's the due diligence that we need to do. And I think I'd first concentrate on the technical due diligence on these aspects, and I I'm happy to to help with that on on on on the mailing list. And then the process things, which working group, and so on. I'll I'll leave that. [01:22:04] **Mike McBride**: Yeah. That'll People with more administrative knowledge, let me Sure. And that'll take care of itself, probably. Thank you. Good comments. [01:22:12] **Tony Przygienda**: It was a tall guy under Mike. Hey. So Tony p, HP Juniper. Thanks. Very balanced view of the world. Some interesting new information, right, where these AI guys are driving multicast. First, one, technical corrections. So beer is bidirectional, the tree kind of thingy. Right? So it's not the unidirectional. The other one is that we have 64 five k receivers, but we have subdomains. So you could stretch it to a couple 100 k. Right? You talk millions, the game starts to become difficult. Right? So but that's just mild technical, you know, aspect. I think the layering is a little bit modeled for my taste. I think you point out certain things where the header would have to be modified. So congestion control. Right? You don't want to push it behind the header. It's really like a header issue. The egg game, I would argue, don't want to push the stuff into the header because you fought the PFE update stuff, right, which is the classical, reliable multicast problems. You have no choice. And it depends very much what the design space is. And very often, you end up with things like negative act designs and things like that. Right? So you design a specific act thing compression, and it may not work at all. Right? Or you end up with 700 flavors. So I think that's something that goes behind the header and becomes this thing that runs on top. Right, and somehow delivers reliability in the desired reliability. I think that in network computation is exactly the same bucket. Right? That's probably not something you want to push into the header. It can go behind. So I think that's the layering thing. And if you look at this egg implosion flavor, for example, the OEM, which we're checking for beer, right, as the beer ping has the same kind of problems. Right? And every time the flavor is different. There's other interesting problem. For example, beer looks pretty cool because every time you can address someone else, it turns out that depending whom you address, your replication will be pretty crappy because the replication points are sometimes a very good idea. Right? The receivers are at the very end, the last replication guy, so you save a lot of volume. But if the replication point for all the receivers are all over the place, you end up almost with Unicast replication if the replication points are very close to the source. Right? And those are problems that also needs to be chatted through for the whole thing to work well. Right? So there's there's more stuff, but I think what you threw up there is is excellent. It's very, very well balanced. It's a good gathering point because the side meetings are, to put it politely, stream of consciousness, to put a little bit negatively, just a zoo. Right? I mean, how do you record and compile it without something like that? So, yeah, thanks. I mean, like, push in this direction and see what what comes out of there. Yeah. [01:25:09] **Mike McBride**: Thank Thank you. We're trying. Yep. [01:25:13] **Nokia Representative**: Yeah. I'm with calling Nokia. I I can't believe I'm agreeing with the Torless for once. But alright. I'll put this in my calendar. I I think with the the act aggregation, I I agree with him with this application, so we might have to put that on the side and try to actually understand what multicast will bring into the picture for the ACK aggregation. If you put that into the side, then I think the picture becomes a little bit more honestly, I cannot understand what they mean by ACK aggregation. Like, I would imagine each one of those GPUs would reply at a different time. So if you wanna aggregate it, there is a bunch of buffering and everything that you need to put in there. So I I would imagine that that would make everything much, much slower if you go to a huge scale. So I think there's a lot to understand there. So it might have a complexity that we just wanna put it for to the site for now and concentrate on the from the LLM server to the to the GPUs to figure that part out. That that's all. Yeah. [01:26:26] **Mike McBride**: That's a good comment. Thank you. Thank you. Yeah. [01:26:29] **Stig Venaas**: I want one more. Yeah. [01:26:34] **Jeffrey Zhang**: Okay. Jeffrey Dong, HP. Peer extensions to address these problems. I actually have some good thoughts that I was not ready to to to present in this IGF meeting. I should have something to to share for the for the next one. All those related gaps out there, I think I have some thoughts on that. So it's good. I think we can do some good PR work in its current architecture without too big changes in the the PR architecture. [01:27:15] **Mike McBride**: Okay. Thank you. Alright. We need to move on. Alright. How are you? And we'll just need to kinda keep it a little brief. That went a little bit longer. [01:27:31] **Futurewei Representative**: Good morning, everyone. This is from Futurewei. This talk is kind of relate to the previous talk Mike just gave. As he said, a point to multipoint pattern is a dominant pattern in AI networking. We need to also realize the reverse multipoint-to-point pattern is always come in hand in hand with a point to multipoint. There are several use case to prove this. We all know that collect collective communication is a key bottleneck for large scale AI training and inference in data centers. Some of them involve all-reduce MOE dispatch and the combined and the reliable multicast that we just discussed or involves a network aggregation process, which engages the reach to aggregate data in flight. It has many benefits, like reduce traffic, reduce packet latency, and also free the host cycles. And, logically, it forms a overlay tree on DC fabric. You can observe this exactly in the opposite direction of the multicast tree. They can have a different shape, but they share this common root and the common set of leaves. So this draft, we will provide a folding layer framework, which use a unified encoding for both trees. We have three first cast use cases or reduce MOE and the reliable multicast. So this slide shows a figure. So you we can see the from up up direction as aggregation, and the reverse direction is a multicast. So we can use a encode, give each node, worker node, an index. Therefore, we can represent the worker set as a bitmap. If it's in the really sparse mode, we can just list their index as a set. So since every point can get index, so we can efficiently to implement the aggregation tree from the leaf to root and the multicast the tree from the root to leaf. So to represent aggregation tree, we simply need to configure each node is which is a switch in the network fabric to tell them which set of workers you should aggregate. And, of course, each packet will carry a header to tell, okay, what's a worker ID. And when it's aggregated, header will be update updated to tell currently what has been already up aggregated and continue work up to the tree until it reached the root. All the all the things has been done there. So, reversely, the multicast tree, we can see is very similar to the principle of the beer. We can just simply use a beer like stateless forwarding. So the two opposite directions can use the same set of encoding method, although they can have the different shape. So the benefit of this unified encoding is the first of is very compact. We can use a a simple nodalist or the simple bitmap to do that. We can support state list multicast. And for the aggregation tree, it's self describing. And, also, it's a topology independent because now we only encode the endpoint. So it's in regardless what's the tree shape of the topology physical topologies look like, it can change. It doesn't change the end endpoint node set. And if we use a bitmapping coding, it's a hardware friendly because we only use simple logical operations to do the to do it to do the to decode the header and perform the aggregation or multicast. And it supports the highly dynamic scenario. It allows a single single package to use a single in different tree, So it's very flexible. And, also, it's a direction agnostic to allow us to use one tree to drive both directions. So why not beer? As we see actually, at least for the multicast direction, we can adopt beer's bit position encoding and you reuse a similar replication behavior. But beer itself cannot realize framework. First reason is that we have a direction here. We have aggregation direction and the the reverse multicast direction. So at least we we use the same encoding, but we need at least some information to tell what should we do is for for to aggregation or to multicast. And, also, for the AI particular job, we needed to, distinguish the different job or different tenant ID. And also for the flow packet, maybe each packet will takes a different, tree than we need in order for the correct aggregation, we need to have something like sequence number. And, also, for the scalability, we can have support different type of header encoding style, like either bitmap in the dense mode or the node list in the sparse mode. So that means we need support different different encoding style than which beer right now cannot support. So here set us some thought about how we support a really huge multicast tree or aggregation tree in a different mode, like a dense mode or sparse mode. For very large possible set of leaves, we can have a two master to to make it a scale. First, we can horizontally partition the partition the tree. So we can have the hierarchical bitmap. For example, at the each a port level, we have a port level bitmap. But when it's go up, we can summarize all of this into a single beat to the upper upper level. It's very naturally to fit in the data center topology and can reduce the header overhead to the most extent. And we can also do the vertical partition to partition the topology into a multiple separate parallel trees. So each trees can have a much smaller bitmap to encode the leaf nodes. And in another use case like the I MOE, for each multicast or aggregation, we call that a dispatch and the combined, it only involve a very small number of workers. But the total possible set could be large. So in this case, if we still use a bitmap, it can be very efficient. So, therefore, we can just use a node set node list to just put their index in the list to encode the set of workers, active workers. But by doing this, we can we can still follow in the similar forwarding behavior. There's nothing changed, but it's just the different encoding process. So as a trade off, we can in a single node, we can support for that bitmap, hierarchical bitmap, or not the least. We can even combine all of this as three in a single solution. It's very flexible. So the status of draft is we in this draft, we just provide this framework and the three first class use cases and give the gap analysis, some analysis on scale and encoding, and some security considerations. So for the group, I'd like to ask the audience to consider, is this encoding framework a useful basis for in that for aggregation standardization? And maybe consider the preferred home for this. Should we consider to extend beer or some new encapsulation protocol should be defined? And, also, we we'll we'd like to see the interesting review and the collaborations. And anybody would like to work with us, please talk to me. And, of course, any feedback are welcome on the list. Thank you very much. [01:36:51] **Tony Przygienda**: Yep. Tony p, HP Juniper. So for reasons unknown to golden man, you're not presenting them in beer. So I've shown up here. It's a mixture between misunderstanding how beer works and reinventing. So the direction is actually I think you are misguiding your understanding of beer. If you set the bit towards the root, you will naturally get the other direction. Beer is bidirectional by nature. It's a it's actually a PIMC, which includes even selective at the same time and directional. [01:37:25] **Futurewei Representative**: It's used for multicast. Right? A hair I mean But you can from send any [01:37:31] **Tony Przygienda**: to network any set of destinations. Yes. There is no direction. So you can send down, and then if a guy sets just the receiver beat to back, then the thing will go back on the same tree. [01:37:40] **Futurewei Representative**: Oh, yes. But still the semantics is a tool. [01:37:44] **Tony Przygienda**: Is no semantics of direction in beer. Beer is a PIMC, which is selective and inclusive and bidirectional. I know it's confusing. Alright? So you don't need a sense of direction. The other things that you said, like [01:37:56] **Stig Venaas**: No. [01:37:56] **Futurewei Representative**: No. No. I I mean, the direction here means the semantics for aggregation or for multicast is two different operations. [01:38:03] **Tony Przygienda**: But that's what you put [01:38:04] **Futurewei Representative**: in different node has two distributors. Well, the [01:38:07] **Tony Przygienda**: root doesn't say x. So if you see an act, you aggregate. Right? So the act tells you semantically that's what you aggregate, and the other stuff, you multicast to everyone. Right? And you can even go and look at the act because the act has a single beat that it sets. So you can just look at the beer head and say one beat set. It must be an act. Otherwise, you do something funny. So I see I think that's, like, a major understanding how to be a replication thing works. The sparse set is an interesting discussion. There is already something in beer, which you probably didn't see, which is called u beer, which is under discussion hanging there. Okay? So if you're interested in that, it does literally that. The this this lock bitmask is an interesting proposal. Don't know I as lock to end bits. Yeah. So the node list is already proposed. Right? The hierarchical bitmap is an interesting problem. I don't know how do you want it to work. Which one? The hierarchical bitmap is an interesting proposal. I would like to see the details how that's supposed to work because once things get big, you get into sets. So this thing, yeah, you could do a new thing if the group is really small, really stable. You could do clever stuff. Yeah. But if you start to go and move towards solving basically the large number of receivers, few trees, then you're basically reinventing beer with a sparse set. The hierarchical thing is the only interesting new thing that I'm seeing. All the rest, like I said, is either already been done in beer or otherwise, there's a misunderstanding how beer works. And the things that you pointed out, I think the previous set, which was that you need transaction ID or something Sorry. The previous one doesn't flip. Right. So the job tenant ID sequence number, that's the thing that go basically carried within the multicast. I don't think you want really want to stash it into the header because then you basically build the application specific multicast. Right? We originally were supposed only to work for l three VPN, EVPN, and encode all [01:40:09] **Stig Venaas**: that Yeah. [01:40:09] **Tony Przygienda**: Files right there, which we threw out. Right? [01:40:11] **Futurewei Representative**: Yeah. So because we want to combine the in network aggregation and the multicast, so many of this requirement actually come with [01:40:19] **Stig Venaas**: Which which [01:40:19] **Futurewei Representative**: has in network aggregation. [01:40:21] **Tony Przygienda**: Is a really poor idea. You can smash all the layers together and have one single bit, meaning everything, and it ends up in tears because you're basically building a failure domain where one thing brings everything down. [01:40:35] **Yisong Liu**: But, you know [01:40:35] **Futurewei Representative**: Media has already realized in that reputation many years ago. [01:40:40] **Tony Przygienda**: I know how Sharp works. Yeah. I know how Sharp works. [01:40:43] **Futurewei Representative**: Noted new idea, and we just now want to combine this with a more efficient encoding to support both [01:40:53] **Tony Przygienda**: But for example, something like smashing sharp into a single head over the multicast with the beer, I think, is just a poor idea, but you'll find out. Okay. So like I said, most of the stuff has already been done in beer. The hierarchical thing is interesting and the direction, I think we have to talk how beer really works because it's basically already inherent in the switching metric with beers. Thanks. [01:41:14] **Mike McBride**: Just a couple more minutes. We gotta get going. We got one more presentation. [01:41:18] **Jeffrey Zhang**: Jeffrey Zhang, HPE Juniper. Can you go to the next slide? Yeah. I just want to add to what Tony said. Pretty much if I understand this correctly, all these things on these slides, beer can already do or we have already extension proposed. For example, multiple parallel trees. I believe it's basically means the multiple sets in BIER. Maybe maybe not that's not the case. We we can talk more offline. Hierarchical bitmap, maybe it's just the the segmentation. For example, you have a large domain. You separate it into different sub smaller domains, and every border the the border routers become a a BFER, and then and so you basically domain to domain. Maybe that's not what you mean, but we can talk offline. But if that's what you mean, it's already supported. And sparse sets, no list, as Tony said, is already been proposed. We have not pushed it forward yet. And and then another thing is that we you are using it in the MOE environments, and most likely, the the the source GPU, source server can choose set as a set of experts that are in the same cluster so that even you you you do not even need to to use the node list. You just when you have a they are if they are together, most likely, you can just use one set to send traffic there. So I we can follow-up offline, but I believe all these things can be can be done already in beer. Thanks. [01:43:03] **Toerless Eckert**: Yeah. So the the sparse sets and the hierarchical bitmap, that is pretty much exactly what we implemented with RBS, which is the first draft in 2022. In the meantime, when when the professors looked at it, we also came up with the combined, you know, hierarchical bit sets and or list of receiver identifiers. Right? And that's all based on your representing the tree through with the which it goes through the network. There have also been kind of almost as as early as beer itself, been proposals to have these hierarchical bit strings explicitly for well known data center topologies from research organizations. Right? So they don't adopt themselves nicely to arbitrary topologies. But if something needs to be optimized for one particular kind of we built a couple 100 of billions of dollars with the same data center type, those are optimizations that have already been defined. Right? So I think all these things are well known. The hard work is really just to figure out what exactly are the use case requirements. Can we do it with, you know, any of the tricks in beer? If not, then we have the proof point of do something better, and then we have a rich, you know, set of preexisting things. So we don't need, you know, any more groundbreaking new ideas. Those those have already been theoretically and then, you know, prototypes been done. It's now the hard work to figure out matching, you know, the business cases with with the solution framework. [01:44:38] **Mike McBride**: Thank you. Oh, you're good. I appreciate it. Yep. You got your fifteen minutes. [01:44:44] **Prasad Miriyala**: I asked for thirty minutes. Got fifteen minutes. I I'm going to speak with double speed. [01:44:50] **Mike McBride**: Put the microphone a little closer to you. Yeah. You always do [01:44:53] **Luis Contreras**: so. So [01:45:00] **Prasad Miriyala**: alright. Few things which we are going to talk about is mostly with respect to PIM. How can we simplify it further for operators? So now I think most of the time, we think as engineer, we come up with a very complex solution. I'm trying to see that can we look it from the other side? Who is the actual user of the PIM? So there are three ideas or rather three problem statements which I'm going to talk about. And the reason why I have not put any solution because I have some solution in mind, but I want to hear from others as well so that we can come up up with something which is really useful for operators. The first one which I'm going to talk about is source discovery. ASM. ASM has been a pain for everyone for last many years. And you talk to any customers, operators, even vendors, everyone does agree that there is a problem. All the state machine which has been defined in PIM protocol, it it is pretty complex. So when I say complex, it is mostly related to when you have to debug something. When you have to as a operator, if you don't really understand complete PIM state machine, it becomes very painful to see where exactly traffic is going. We did work in this area. We came up with PFMSD, which was developed in same working group. We have implementation, but it has challenge with respect to scaling. So if you have large network and if you keep keep flooding your source information continuously in the whole domain, that is not something really customers or operators are expecting to happen. The second thing is if you don't have the FMSD, then you have to live with RP. To have RP, you have multiple protocols, whether it is static RP or bunch of dynamic RP protocols. Then if you have multiple domains, you have to run MSDP between them. So with PIM, actually, actually, we are adding five more protocols on top that makes overall experience very hard. And when we are migrating to I p v six based backbone and coming up with PIM based solution, again, these problem statements will continue. So I started thinking about it that how can we get rid of this whole RP business in the backbone. We cannot get rid of ASM, and the reason is very simple that there are millions of set top boxes which can do only Star,G join today. And none of the operator will bear the cost. Me me as an end user, I'm not going to update my set top box because my operator wants to do IGMPv3. So that that is not possible. So we'll have to solve it in the backbone. One of the probable solution we were I was thinking of is coming up with something lightweight similar to be a route reflector. And I was talking to Jeffrey. Thanks, Jeffrey, for comment that maybe we should consider using BGP itself rather than coming up with something new because BGP has been well defined protocol, and it already knows how to scale. Hierarch you have to do if hierarchy if you have multiple RRs, how they talk to each other. There is a well defined process. So maybe it will be worth considering that can we announce a source BGP itself? So there is no more shared tree. You directly learn about all the sources. You do direct SD join. That is one of the problem statement. Though if we had whole thirty minutes, we could have spoken for long. Do Do you you want want me to cover all and then ask question? I see you in queue. [01:48:41] **Tony Przygienda**: Is it sure? [01:48:42] **Toerless Eckert**: So so you you you you're you're you're saying just, you know, use BGP to announce information for SSM mapping? [01:48:49] **Prasad Miriyala**: Kind of. Yes. We'll go to second problem statement, which is MOFRR. So today, seventy four thirty one defines the base MoFRR, and I think all of us majority of the vendors have implemented it, and it is being used as well in the field. There are two types of MoFRR which is possible today. One is RIP based, and second one is flow based. So let's understand what happens in flow based MoFRR. In the last half router, if you have network something like this, in the last half router, you can have some kind of monitoring here where you are really monitoring both flows from the red sorry. Not red. Blue and green path. And you are accepting one of the traffic. If there is a failure happening anywhere in the network, you will be able to detect it fast enough. And once you detect it, you immediately do this RPF switchover. Life is very simple from outside. But implementing this, it it's very hard. So when I say implementing this very hard, what it means that ASIC really takes a toll. So if you start watching or monitoring each flows, something else had to be dropped off. So, basically, the load on ASIC becomes very extensive. So it becomes a costly solution. The second one which is used mostly is RIP based. So in case of RIP based, from last hop router, I have two next hop. It could be CMP or U CMP. And when I start sending the join, join the primary path and backup backup path. And once we start getting the traffic, you will have based on some logic, it could be any hashing or it could be priority. You accept one of the path traffic, and the second second one, you start dropping it. So you have to really detect that something failed before you can move to the backup path. So let's see what different type of failures are possible and how we react today. So if there is a failure here, LHR can immediately get to know that there is a failure, and it can switch the RPF. Life is simple, fast enough. You won't see the glitch. What happens if traffic my this link fails, which is FHR to p one? Again, same thing. LHR will be able to detect the rib rib next hop change. And right now, even though I have made as a blue and green path, don't consider it as a flexalgo. You can consider it as a still two path. So the rib next hop changes, I'll be able to switch to next path immediately. But what will happen if my sorry. If failure is between p one and p two. So in case of p one and p two failure, for LHR, there is no next hop change. LHR is still pointing to p three. P three is going to take care of the local repair. So your team join will take care of going to the different path locally. What this result into that you will have a traffic loss in the receiver. So receiver sees the traffic loss. Last Last half router has both copy, but it's still it cannot do anything. So this is the problem statement which I'm trying to [01:52:02] **Nokia Representative**: talk about. It's too much. Just for everybody to understand. Today, there is a solution for that. It's not optimized, but there is a solution for that because p one sends the traffic active to p two, but then there's gonna be a backup traffic to p four to p three where p three drops it. Right? But what you're kind of saying is that that's sending traffic to every single link, and it's congesting the net far. You're trying to optimize that. [01:52:30] **Futurewei Representative**: That's correct. Okay. Yeah. Just to make sure. [01:52:32] **Prasad Miriyala**: Yeah. That that was my next statement anyway I was coming to. So today, the way we solve this is you turn on everywhere. So if your network is large enough, every hop, you'll have to create two joins. So depending on how many hops you have from the CO to source, that many copies of the traffic will be flooded across the network. So what we are trying to do that can we can we do some more simplification here? Or is there any way to really optimize this overall solution? One of the solution which is in mind, and right now, it is not updated in the draft. Let's say that what exactly we are trying to achieve. We are trying to achieve two things. In the last half router, I should be able to detect the failure anywhere in the network. The second part I want to do is I should be able to that this network failure has impact on my what all trees, three one two three, whichever trees. And then based on these two parameters, I want to make a decision. So one of the solution which comes into the mind is detecting detecting the failure. Today, IGP already distributes your link state. So any change in the link state, we already know. Last half router does get the notification. So we have that information already in the last half router. So one part is solved. We have we have this information. The second part is how will I know whether this particular failure really is going to impact any of my tree? And this is where we need some kind of modification. Again, keeping as a engineer side, rather I'm thinking as a operator. I don't want to overload TIM protocol to give me this info that where exactly my tree is going from. And one of the probable solution which we were thinking of or I was thinking of is can we use telemetry where today telemetry anyway exports huge amount of or many vendors, they have implementation where you can export your SG info that each tree is going from where, which link and which router. And if we are able to also export the IGP unique ID associated with the link along with SG, somewhere, it could be in the LHR directly or it could be some external you can consider at a controller is the wrong word, but external processing center where we are able to get the whole tree view, then last half router can really merge both of them together and take the action. Tony? [01:55:03] **Tony Przygienda**: So sound sound like rust white. So it's the p etymology gets in the way. Okay. What does it mean, really? What makes you think that any mechanism, when you're looking a couple hops away, will be faster than the fifty millisecond local repair there. See whether it's a so let's say you you start on the LHR to look at the p four UK using some network telemetry. Okay? If the p so let's no. If the p three creeps out, you're totally done. But let's say the the the p four telemetry path goes some way and something in the middle fails, right, which has impact on the multicast tree, but the telemetry also has to reroute around it, and that needs time. And the farther you look back, right, whether you use telemetry, you use what makes you think IGP flooding will be faster than local per fifty millisecond before IGP coughs up all this stuff and then coughs up the SPF, and then you'll be probably at the same same point. The problem is that you want to look far ahead, but you need some time to diffuse the information unless you entangle the bloody things. But then it's a different working group. Right? If you don't have entanglement, all these things have delay in propagation. There is some kind of, you know, diffusion front, and there is no magic to shorten this diffusion front. Right? So I I don't think there is much you can do here from the very foundation of the problem, but, you know, tell me. Mhmm. [01:56:37] **Toerless Eckert**: So I remember this discussion from 2008 with Comcast. Right? So it's kind of totally annoying that seemingly none of the vendors has managed to do, you know, packet reception based MOFR failover as opposed to control plane based failover. Right? And I so, I mean, that that was the whole point of doing, you know, MOFR in the first place that we can do a traffic absence based failovers. There are also products that do this seemingly not any of the the call routers. Right? I remember several of the video products that that are supporting that. So if I may be very self serving, I would say that it's perfectly fine that vendors that have these type of products suffer the pain. And if you wanna see a better solution, I'm I'm I'm recommending to, you know, listen to, BIER-FRR in in the afternoon. Right? So that provides a perfectly well working sub fifty millisecond solution failover. [01:57:40] **Nils Warnke**: So Neil Sank, Deutsche Telekom. This failover scenarios to me are very real because we experience them, and we have a lot of pain in the past. And this fifty millisecond failover promises have rarely been kept. Let's put it this way. What I can spill is that as soon as you are above one fifty milliseconds failover, you experience some issues in the picture with TV products. As soon as you're above three hundred or two fifty milliseconds, you experience a still picture or an error picture. So these fast detections to us is really a key issue and a key thing that we're looking for. What what I've seen is, especially, RIP tends to be rather slow. So any optimization we can find there is really appreciated. From a flow based approach, we've looked at this, and this is something that we've really grasped for since long. But knowing larger deployments and how heavy this is on the NPUs and the hardware inside the routers, this has not really been something that scales. So I'm really curious on what we can optimize there. I'm not quite sure, but if we can leverage messages inside the routing protocol, for example, where we can either up the tree or down the tree message these kind of failures and give an give an indicator to to fail over quickly, That's been a discussion that we've been having a couple of times as well. But honestly speaking, not being an expert in these route route messaging advertisements, I can't really judge if this is an option that's feasible for looking into. Sure. [01:59:57] **Sunny Zhang**: Sunny's on the key. In our implementation, we borrowed some method from BFD detection for the fader fast detection, and we do some special special execution on the remote detection, such as p one can report some key fader events to our HR. So our HR can change the way quickly. So we I think, in my opinion, it's interesting topic, and we can talk it together and to help to to it's better. Yes. [02:00:34] **Prasad Miriyala**: Sure. So I'll I'll answer the Tony's question first. [02:00:37] **Mike McBride**: And then we gotta kinda wrap it up. I'm sorry. [02:00:39] **Sunny Zhang**: BSD. Yeah. [02:00:42] **Prasad Miriyala**: Yeah. So about your question of I got your point. So I'm answering Tony's question where the fifty millisecond local repair. For PIM, if there is a failure between p one and p two, p three has to do a bunch of things. First, detect it. Second, Second, do the RPF change. Third, program the hardware. Right? And then you send the join via this path. Depending on how long this path is in this picture, it is very simple, just one hop away, where it could be multi hop. The time, you cannot guarantee that it will really recover in fifty milliseconds. So that is number one. The second thing, telemetry is just one time in advance to understand the tree. Telemetry is not being used for all the signaling. The signaling, what I'm using is your update from the IGP. So IGP update, which you are flooding, definitely, it is going to reach much faster than much faster. I won't make it a generic statement. But in the high scale, yes, it will be much faster for LHR to understand that that link failed compared to this whole processing happening in these core router. With the PIM signaling. Yes. [02:01:47] **Tony Przygienda**: Yeah. I can yeah. I I agree. [02:01:50] **Prasad Miriyala**: About BFD, BFD is definitely one of the choice. But per tree, tree, you have to maintain a BFD and then act on that. Yes. So it it is not going to scale. [02:02:01] **Luis Contreras**: Yeah. [02:02:03] **Prasad Miriyala**: And thank you. [02:02:04] **Mike McBride**: Thank you for your time. Sorry. I didn't get it. Yeah. We've got one more. [02:02:08] **Prasad Miriyala**: Yeah. Yeah. I saw what Yeah. [02:02:09] **Mike McBride**: We don't have yeah. Sorry. Thank you. Next time. [02:02:12] **Stig Venaas**: Alright. Do you wanna Yeah. [02:02:19] **Mike McBride**: Thank you so much for your time. We'll continue a lot of this on the list. Let's go. [02:02:25] **Stig Venaas**: I kinda wonder if this point to multipart DFT might help in another way. I think it's, one DFT session for the whole tree or something. [02:02:35] **Mike McBride**: I don't know. But I I wish we had more time on that. That would have been better. Yeah. [02:02:41] **Stig Venaas**: It's an interesting discussion. Think so. Yeah. Yeah. I know we usually end up having too many presentations, so we want no discussion.