Markdown Version

Session Date/Time: 23 Jul 2026 09:30

[00:00:29] Carlos Bernardos: Okay, guys. Let's get us started. Please, try to close the door. It's a bit tough, but even if you oh, okay. Okay. Thanks anyway. So this is a FAN working group. I'm Carlos Bernardos. I'm the co chair, Eddie, run this online. This is the first session of this working group. So welcome everybody. Please, if if somebody remotely attending can confirm me on the chat that the audio is okay, that will be great. Let's start with the note well. It's already Thursday, so probably you are all familiar with that, but it's important. So if you are not familiar, please take your time to scan the QR code and follow and read and follow the instructions and guidelines. Because if you participate in the ATF, you have to follow the rules. There are two points that are very important. The first one is we have to be professional and polite and respectful with our colleagues and the way we communicate comments and disagreements, if any. And, of course, if you contribute, you also have to abide by the IPR rules of the ATF. Some tips or guidelines important if you are here attending on-site. It's important that you also register through the Mitico tool, probably using the light version because this is what is used by the secretary to know how many people attend the meetings and properly dimension the next ones. And if you are remote, you should also use it, of course. And in if you are using the full client version, please try to keep audio muted. We have minute taker. There are there is a tool online, so we all collaboratively can take minutes. That will be appreciated if somebody else in addition to Mike. By the way, Mike, thanks a lot for taking the notes. Can join there and help Mike take in the some notes. As usual, the important thing is the agreement or significant comments. We don't need to transcript. There is already a transcript too, so we don't need that. Some resources for this meeting, the agenda, and the usual stuff. I will not spend more time on this. So this is the actual agenda. We will well, we are starting with this first introduction about the administrative staff and a bit trying to summarize the charter of the working group. Then we will go into the product statement requirements and gap analysis, which is the main discussion topic for today. Then we will have a presentation by an operator on some of the practices that they are using and how they see potential future work in the interconnection case. Then we have some additional presentations on use cases and gap analysis on the different scenarios that are covered by the charter. Time permitting, we will have a presentation about the frame the fund framework related to the fund framework. That's time permitting because basically, the goal for today's session is to focus on the first milestone of the working group, which is what we have chartered to do at the end of August this year. So a quick summary of Charter scope and and deliverables. Our mission is to convey locally detected network conditions to remote nodes. Two, with the aim of enable more efficient and responsive traffic flow handling and develop a comprehensive solution. We have a couple of kind of priority deployment scenarios. That doesn't mean that we are constrained by only those two, but those are the ones that we should focus on at the beginning at least, the data center and the data center interconnect. We are also or the charter provides examples of what we should consider as the notification triggers, the inferior signal degradation errors, output queue congestion, but it could be extended later. And the work plan is to work first on the product space and do only solution work after working group last call on requirements and framework. So that's the the idea. In terms of deliverables, we have a product statement requirements and gap analysis, which is an informational document not aimed for publication. So the idea is to adopt a document, and that's the main focus of today's session to keep it as the baseline for the additional work, but not to be at the end of the process published as an RFC. Then we have the framework document describing the architecture and integration with existing mechanisms. Then the protocol specifications. By the way, in the charter, we say that we need at least two interoperable implementations of those protocol specifications. So that's an important thing. And then applicability and operational guidance support and young models for the configuration and management. So those are the deliverables. Bless you, team. And those are the milestones. So as you can see, the first milestone is in August, the to adopt the problem statement and requirements and gap analysis. Our intention at CERSE is to run our working group adoption call on the document that is gonna be presented that the charter said is based on a document that was already adopted in the routing working group. So, basically, we ask the authors to submit that as an individual submission on this working group, and we will run, follow the process, and adoption call on that one. Hopefully, there will be no major issue. And, of course, the working group the document will belong to the working group at that point if adopted, and we can address additional comments, suggestions, and changes. And then in November, we have the framework document. So, hopefully, if we can adopt the document, the first milestone right soon after this meeting, we will start working on the framework document. And the idea that the by the November meeting, we can adopt a framework document baseline document draft to work on that. And then the rest of the the milestones will fall. So as I mentioned, the scope for today's session is the first milestone. SCOM is product statement requirements gap analysis against the the 16 scenarios, the DC and the DCI. And we it's a lot of scope for today. Please, if you can close the door, please. The ones that sorry. Okay. Out of scope is to any kind of protocol solution at this point because we are far from that. We have to adopt first the framework document after the problem statement and gap analysis.

[00:07:34] Tim Chown: And with that, we can move to the

[00:07:38] Carlos Bernardos: next presentation. I think it's you. Yeah.

[00:07:58] Jidong: Okay. Good morning, everyone. This is Jidong from Huawei. I'm going to give a presentation on this, fast network notification problem statement on behalf of my on my co authors. So the first part is a background. I think, for people who have joined the discussion in the FAN BoF and also the previous discussion in the RTGWG, this should be very familiar. Now we'll go through it quickly. Basically, the motivation is that, we see that the modern network applications, especially the AI or machine learning service and other cloud services require the network to be adaptive to the presence of the failures or the degradation or congestions. And currently, the existing mechanisms in the routing or traffic management, often face limitations in either the responsiveness or the coverage coverage and operational complexity, especially in the large scale and the high bandwidth networks, which is the DC and inter DC scenarios. And, in these cases, fast network notification can signal the network conditions, include the adverse condition and their recovery to relevant nodes in a timely manner so that they can take fast and efficient actions to the quick critical events so that we improve the service quality and also the network utilize utilization. Basically, this document described the problem statement and the gap analysis of the fast network notifications. It was used as a major supporting document in the incubation of this working group in the RTGWG. So here, we will take a look at a problem in an example, which is the the inter data center AI training cluster case. In this case, we know that the AI training service require ultra low latency and high throughput. Otherwise, there will be a failure or the lost or the waste of the resource. In this case, we see that in case of failure in the fiber, they can disrupt the entire training job, and the delay in the recovery can cause a waste of the computing resource, the energy, and also the time. With the existing mechanisms, we can see that we can do the detection with BFD for the link failure. And, also, we can take the fast reroute, which will find a backup local repair pass to steer the traffic away, then the road can convergence can happen based on the update of the topology information. But the problem with this existing mechanism is that, like, the BFD detection will take several tens of milliseconds and the fast rerouting, although it can find the alternative pass, it is not aware of the quality of the other pass and may cause congestion on the backup pass, maybe can impact other services as well. And the route delay in the road convergence are not acceptable to this kind of AI training services. Here, we also summarize the problems and gaps in the existing mechanisms in this draft. Basically, we list four major problems to address. The first is the slow dissemination, which just means that the existing mechanisms, either they use control plane, hop by hop distribution of this information is considered relatively slow for these requirements. And, also, for the existing notification mechanisms, which take the signal to the receiver and back to the sender, we'll require a round trip time delay. And this will delay the reaction to this network conditions, like the failure or congestion. So what is missing in the current solutions or mechanism is the faster delivery of this network notification information, preferably in the data plane. The second problem is that the cost grained signals in the messages. Currently, most of the existing magnum would take a carry only the binary or threshold based information in the message, which is not enough to take fine granular or precise actions by the notified nodes. And what is missing in this is we need a more fine-grained information in the notifications. The third problem is limited visibility, which is that usually the current local protection behaviors can only have the local view of the network conditions, which is makes them suboptimal to the network. The and some of the optimal decisions may impact the other part of the network and may not be a good way for the global network opt optimization or the other service performance. So it is what is missing is we need a wider view or, like, a better understanding of the network in a larger scope. And the last point is overhead and the scalability. We know that there's existing mechanisms which can provide very frequent on the on high volume information either to the control plane or to the network. But this will cause very large overhead to the nodes to react very quickly so that we can what was missing is light lightweight and also selective notification mechanism so that we can do this reaction in a very fast manner in the hardware or in the data plane. Okay. In this draft, we've also identified the problems which will be need to be considered in this, faster notification, scope. The first is, what kind of information need to be carried in this fast network notifications. And the document has identified several types of informations, including the event type and also the look location of the event, the fine-grained network status information, and the pass identification information, and also the flow or service identification information. And this may be some of them may be needed for sub specific scenarios. Well, we don't expect all this information will be carried in one notification message. And for the recipients of the fast network notifications, draft also lists several types or the groups of the notification recipients, including the agent adjacent nodes, like routers or switches or the non adjacent nodes, which are maybe several hops away from the node which detect these events. And it can also be the ingress routers or switches in the network. And option another option is maybe some mass notification message can also be sent to the edge nodes or the hosts. Network controller is also considered as one optional recipients of this con notification message, although maybe that requirement will be slightly relaxed comparing to the data plane recipients. And the recipients, of this specific notification message may be determined, based on the configuration or the signaling or via some subscription mechanisms. And the next thing is, the delivery mechanism of the faster notifications. According to the type of recipients and the group of recipients, several options can be considered. The first is we can unicast directly a mass notification message to the one of the recipients. And if there's a group of recipients need a same notification, we can also consider to use a multicast. And in some cases, the hop by hop delivery of this notification to a series of nodes along those specified paths is also considered a use case. And flooding in a specified group of the network is something maybe useful if we need this information in the region, which can take the proper actions. And the draft also highlight that we would prefer to reuse existing messaging and transport mechanisms if they are possible for the delivery of the faster notifications. And, if there are still gaps identified, new up, new protocols may also be considered. And although we know that the scope of this working group is on the notification mechanisms, the draft also, as discussed previously in the RTG, we listed the candidate actions to this notification mechanisms, including the recipients can take the following actions when the notification is re received, including the switching all the traffic from pass to other avail available pass, and the recipient can also steer the traffic to alternate pass or links and modify the load balancing ratio among a group of the passes. They can also send a notification further to other recipients. We think the the first notification draft will not focus on how these actions will be done and how we can, but we will need consider the interaction under the between the notification and the actions. I think the how these actions can be indicated in the passage is also something needed for the discussion. So here are the updates of the draft since last, meeting. First, we renamed the draft to fan problem statement because we have this new core working group formed, and the scope of the draft has been extended to cover both the gap analysis and the problem statements. We also have some tax added about the previous work in about the transport con area, like the congestion control work as according to the TSB review comments. The draft also highlights that the delivery of this notification message need to be based on existing mechanism when possible. And, we also added a newer section, the operational consideration section, to talk about the the following considerations, including the manageability, interoperability, coexistence with other existing mechanisms, and observability, partial deployment, and also interact action with other notifications in the network. So for the next steps, I think we are working on the transferring this draft from RTG to this fan working group as a grid between the co working group chairs. And now we request to working group to adopt this document as a base supporting document of the following works in this working group, and we would appreciate further reviews and the contributions to this draft. Okay?

[00:20:00] Carlos Bernardos: Thank you. Any comments or questions? Yeah. Wim?

[00:20:10] Wim Hendrickx: Yeah. Wim Hendrix from Nokia. I think that one of the

[00:20:14] Carlos Bernardos: Can you speak closer to the mic, please?

[00:20:16] Wim Hendrickx: Yeah. Yes. Is it better?

[00:20:18] Carlos Bernardos: Yeah. I think it's a bit better. Yes.

[00:20:20] Wim Hendrickx: So so two things. One is I think we should also look at the security implications of this. Right? Because if you basically send a message from a router to something, then what are you authenticated? How do you all the security aspects are related to that. And then I think we are discussing here, like, gap analysis and problem statement. I know about I for example, in UEC, we are also discussing the BTS, which is back to sender, which is actually very similar to what this is. So I would recommend that we also look at that information or some of the work that is being done here there. Maybe I can ask to facilitate whether they present in this working group what they are doing just to share some information because it's a different standardization activity because the work that they are doing is very similar to what is happening here. And I think it's important to share some information. And if if there is a solution, that we should not duplicate the work if it's already there. Alright.

[00:21:19] Jidong: Yeah. About the security, I think, that was discussed in previous meetings. I think, we have some security considerations in the draft, and people have agreed to maybe focus on the single domain scenario in the beginning so that we can relieve some security concerns on the solutions. But I think if we want to extend it to multi domains, that will be that we need more considerations. But for the other part, I think I cannot get all of your questions. So maybe we can talk offline. Okay. Yes.

[00:21:56] Jeff: Just to answer NVIDIA, a couple of points. One, two wins command BTS has been defined by IEEE. There has been work start about seven, eight years ago. So if anything, we should issue license to IEEE to work on BTS. EAC is not the right. It's standard today. Mean, they're not standard today to begin with, but inside triple a domain. Problem statement, as never before we understand the constraints of application layer here. But say, AI is the focus really. Right? As you put it in the first slide. The transfers are stateful, somewhat comparable to TCP. RDMA opens the QP, which is stateful channel. It has its own telemetry. It has its own telemetry at different levels, transport level, at the semantical level, at the application level. We need to make sure we don't end up in a situation of SDH over ATM over IP, which each layer runs its own faster route. We've seen this. It's horrible. Right? So we absolutely need to make sure to understand the timing of feedback loop and make sure that whatever happens on the switch or router doesn't interwork in strange ways in what application is trying to do. And there are very basic things such as CN marking to more complicated measurements to really PSN measurements on application layer. Right? So all of this need to be well understood before we decide outside of basic encoding, which is obvious. Not related to that, but timing, feedback loop, and what consumer and who is the consumer, what they do with the information. Right? It's fundamental for solution to work. Otherwise, we'll end up in kind of soup of different things.

[00:23:48] Jidong: Okay. Thank you for comments. I will we will consider it after the meeting.

[00:23:54] Carlos Bernardos: Tim? Yep. Go ahead, Tim. You have to to use the tool to join the queue.

[00:24:19] Tim Chown: Yeah. Tim Chan. So I think, yeah, obviously, the milestone is to adopt the problem statement. This is the problem statement, so I think we're adopting. I think it's in good initial shape. I think we've we've had some good comments. It's quite impressive as well that the the working group already has nine documents in submitted in the in the first couple of months. So it's it's it's nice to see it active. Yeah. I think it's it's it's a good good document, and I support adopting it.

[00:24:47] Carlos Bernardos: Okay. So let me raise run a quick, pool of questions. So one is just if you have read this document or the one that was adopted at RTG just to k. I think that's stable now and sufficiently meaningful. So now let me run another one that is basically asking if you think that this document is a good base for adoption. Of course, we will run this on the mailing list, so this is just to take a initial idea, but we will follow the process. Okay. So we'll run the call for adoption. Thank you very much. We move on on the next agenda. Luis?

[00:26:24] Luis Contreras: Thank you. So this is Luis from Telefonica. I will go through through this topic from from the angle of interconnection of network. So, basically, how an ISP basically could leverage on this for, yeah, basically managing the different I mean, different large sources of of of traffic that's been integrated to the network and how to, yeah, basically support this for automating automatizing the the way of taking actions for kind of first response to any issues that could be produced in in interconnection. And and this is somehow coming from from experience that we already have had in integration interconnection scenarios where actions that are taken in in different administrative domains impact the our our own network and can produce problems on it. So the point is that the interconnection domains are typically characterized by either a large number of of links, of interconnection links, or links with of large bandwidth of both of them at the same time. And, also, with the fact that we could have multiple points of of interconnection, so we need to basically, with a single a given single administrative domain, deal with the these different, yeah, points of interconnection producing a large amount of traffic. The now these these links are moving from 100 giga to 400 giga, and and the evolution will come with 800 giga, one dot six terabits, and and so. So, basically, creating a scenarios of a large bandwidth interconnection where any problem, any issue could could, yeah, create serious problems in the internals of the of the ASP. So unplanned situations could be the the reason of, yeah, producing these these changes in the delivery of traffic of the flows, so producing those problems. And evidences of problems could be, of course, any failure on the on the link or but also trends like the sudden increase of traffic or other signals or maybe instability of the links, aspects that could basically rise the the point of the there is an a problem. I'm a potential problem in a link, so I need to react to that so to avoid that we could have problems in interconnection, even internally, but also that we could create problems to the other domains. So the focus here for for this talk is basically in in using the notification to avoid service affection once these changes are unexpected. So we don't have a plan intervention. We basically are we are running the network. We have these operational issues, and then we need to react fast to to that. For plan changes and anticipated changes and so, there is there are other efforts, and we are trying to move other efforts in in other working groups, like in TBR. So what kind of scenarios could be of interest? Essentially, whatever scenario that handles a lot of traffic, and this means basically classical interconnection scenarios or PD in transit, interconnection with other ASPs, and so. Also, the interconnection of of large content providers, CDNs, caches, basically video streaming video streaming, which is the majority of the traffic nowadays. Also, the interconnection of data centers, so basically impacting with the data center gateways where we could have yeah, large number of links and and links of basic high capacity. And in general terms, cloud edge continuous scenarios where we could have we could be interconnecting different clusters of Kubernetes or compute the cloud infrastructure that where we basically cannot control where the workloads are deployed, and and, basically, we don't have any any control of what could be the traffic pattern that could happen in in a given point on time. So the point is the here, what we what we can get from these fast notifications and the idea will be to have this kind of fair response to a problem to to in order to mitigate it and later on to take care of other decisions, of course, but at least to have some kind of control if something happens in an unpredictable manner. And here, I I wrote a couple of examples. So let's consider an interconnection between domain a and domain b. It could be the case that maybe if there is a problem in domain a a, for instance, the the some issues with the with the links, maybe some flapping phone or whatever thing or any sudden increase of traffic, maybe we could be could handle this and trigger any fast notification. So maybe for for using the well known BGP community graceful shutdown in RF defining RFC eighty three thirty six for maybe notifying domain a that, hey. I I will shut down this this interconnection. So somehow trying to get some automatize automatization, let's say, on the handle of interconnection and later on analyzing the problem and so on so far. Another example could be that they maybe there is a failure on the link in between a and b. B could trigger some internal established policy or dynamic established policies, sorry, or dynamic reconfiguration so that we could handle internally the issue of maybe one link from a a link aggregation has been done or has been down or or whatever. So in summary, basically, taking some reaction once we detect something that is not predicted beforehand. And and, yeah, basically, the final message will be with with fast notifications, we could have this first response and, basically, avoiding to to make the problem more serious and impacting the the ISP and also propagate the the failure to others ISPs. And this is basically the my point.

[00:31:55] Jidong: Thank

[00:31:55] Carlos Bernardos: you. You. Any comments or questions? Yeah. I see one coming to Mike. Jeffrey?

[00:32:09] Jeffrey: Jeff, ask, the minute you're saying put anything in the BGP, this is no longer fast. So what are the conversation points this working group should be examining is fast notifications are lovely. We already had, one of the speakers say that BFT itself is too slow. This will be faster. But if all you're doing is being fast to do global, repair of some sort, you're not doing anything useful.

[00:32:35] Luis Contreras: I didn't get well the the comment. Sorry. I couldn't listen well. I'm sure it's just just a comment or a question.

[00:32:42] Carlos Bernardos: No. It was a comment

[00:32:42] Tim Chown: that if you are talking about BGP, it's already

[00:32:44] Carlos Bernardos: I mean, it's not fast. Right? And and So

[00:32:47] Luis Contreras: of of course. But, I mean, the this the the action would depend on the nature of the of the problem. Maybe you are seeing a trend, an increasing trend of traffic. Maybe you have some time for for reacting. I mean, depends on the granularity of action for for sure. If the links has gone down immediately, you you cannot leverage on procedures that take some time. Maybe you can trigger other actions. These were just examples that maybe based on the problem that you identify, you could have different tools for I mean, once you have the trigger, different tools for managing the situation. Like, we would need to explore the what is the proper situation and the timers for that, of course.

[00:33:22] Jeffrey: Agree. And I see that g is the queue behind me. The the point, I guess, I would take forward from this is that there's gonna be different hierarchies of speed that we're looking at, potentially microsecond level, no reaction mechanisms, millisecond level reactions. And we're also looking at mitigations, which may be in those different things. The minute we have any control plane, mechanism, this is gonna be on the second level of mitigation. And that should be a just recognized by the working group that the control plane is the wrong place to solve classes of these fast reactions. It's okay to have that in the toolkit, but, you know, we've already established we know what that's going to do. Good.

[00:34:04] Jidong: From Huawei. Thanks for the presentation. I think this is also a very use useful use case for the fast notification. I just want to check, in this case, in the interconnection case, do you consider the notification will be sent between different administrators or is still within the same domain?

[00:34:27] Luis Contreras: Could be could be a case here. Basically, two examples are not inter domain. Let's say, basically, the domain b is taking the action for their own response to the to the problem. Being in inter domain, things could be implications, of course, security, all this stuff, but also the trustability. But, I mean, the the point the starting point from my point of view would be to to start with the the actions that are taken in their own domain to external actions that you perceive in interconnection.

[00:34:55] Jidong: Okay. Thank you. The second question is, in this case, do you expect the action will be relate to you to the control plane mechanisms always will be react directly in the data plane?

[00:35:09] Luis Contreras: I didn't consider at this stage any kind of solution. So, basically, just considering the problem. Depends again about the the granularity of the timing that we have for solving the problem, what what could be used. So, essentially, this would help us to have a trigger for an immediate actions faster than human reaction. Basically, it's auto automatizing the the intercourse just to mitigating problems.

[00:35:34] Jidong: Okay. Thank you.

[00:35:38] Jeff: I'm just concerned with NVIDIA. Great use case, and I think this should drive us to think about the encoding. So very fixed, very fast encoding versus something more flexible, potentially TLV based, where we can encode not only potential remote links that is congested going down, but maybe a prefix attached that is under congestion. Perhaps a SEN or some other information could be fed into BGP that BGP can react. Another interesting consideration here is stateful versus stateless solution. So most of solution are draft published as of now are completely stateless. You receive notifications immediately. If you make it stateful, and we say, if we receive number of packets within a particular amount of time, we might trigger some change in behavior. So that should go into consideration as well. Thank you.

[00:36:34] Luis Contreras: Okay. Thank you, Jeff.

[00:36:36] Carlos Bernardos: Thank you. I closed the period just for the sake of time, so let's move to the next presentation. Thank you.

[00:36:55] Xiaomi: Hi, everyone. I'm Xiaomi from ZTE. This presentation is on the problems and the gap analysis for DCI congestion notification. There are three individual jobs involved. Today, I will present only the problem gap analysis, no solutions that will be mentioned. This is the first draft, fast notification packet in RoCEv2 networks. First, what's CNP in RoCEv2? This graph shows the mechanism for the CNP congestion detected by switch, ECM bits in IP header marked by switch, and marked ECM bits detected by the receiver and the CNP sent to sender by receiver. Last step is transmission rate reduced by sender. Note that CNP is already widely deployed in AIDC. This is is the CNP format defined by IBTA, InfiniBand Trade Association. That's an organization outside IETF. This is a UDP packet. There is a UDP destination port already assigned by IANA. And the raising the transport header above UDP, there's a field called the source QP can be used to identify the congested flow. Okay. Then what's fast CNP and why? It's simple to understand that if you have a switch closer to the sender, but there's a long distance between the switch and the receiver, it will take more time for the receiver to give the congestion notification to the sender. Then the switch itself send congestion notification to the sender. So the message is simple that the switch can send fast CNP to the sender directly. It's not needed to send the ECM marked data packet to the receiver first. There is existing implementation of fast CNP, And this graph showed a scenario that's a data center interconnect by OTN device. In the existing implementation, FastCNP has the same format with CNP defined by IBTA. DC gateway needs to populate the source QP into FastCNP, But it's unable to know the sourceQP directly from the congested data packet. So in the current implementation, the g c gateway need to learn the mapping table between the source copy and test copy before the congestion happened. And after the congestion is detected, this gateway will look up the mapping table and find the source code p and construct the fastest CNP, send it to the sender directly. So what's the remaining problems of fast CNP? First problem is that if it's not the DC gateway, but to the spine switch get congested. How can the spine switch to learn the table between the source copy and descriptor before the congestion is contacted it is detected. Because of the traffic, the during the RDMA session establishment, not the forward pass and the reverse pass traffic can go to the same spine switch. So the spine switch has no method to learn the mapping table before the congestion is detected. So that's the first problem.

[00:42:14] Carlos Bernardos: So sorry to interrupt. We have, like, five a bit less than a few minutes, and you should allow for questions sometimes. So if you can speed up a bit. Okay?

[00:42:22] Jidong: Okay.

[00:42:25] Xiaomi: Second problem is that when there are two DC gateways for load balance, how can the DC gateway learn the mapping table before? Because the forward pass and the reverse pass may go to different DC gateway. Okay. This is the second draft. So this scenario is a bit different with the previous one. In this deployment scenario, the data center will be interconnected by IP WAN. It's not connected by OTN device. So the problem is that if the p one node within the IP one get congested, how can the p one node send fast CMP to the sender directly? Because of the p one node is in IP one domain and the sender is within the data center. It's a server. How can the p one node send direct fast CMP to the sender? That's a problem. This is the third draft. It's about the congestion notification for pause. That's a different area than passive CMP. So what's PFC? PFC is priority based flow control, provides a link level flow control mechanism that can be controlled independently for each class of service. Note that PFC is already widely used in AIDC. It's a mature technology and also standardized decades ago. This is a PFC frame format defined by IEEE eight zero two dot one. That's a Ethernet frame, uses to pause the traffic forwarding. So what's the problem? Problem is that if data center is connected by IP WAN in the remote data center. There's a PFC between the data center switch from leaf to spine to DC gateway and then to p e two. What we can do to mitigate the congestion? So that's the problem. The p e two can also can do the first thing is to buffer the traffic, pause the traffic forwarding. But if the PE two get congested too, how can we mitigate the miscongestion? Yeah, that's it.

[00:46:00] Carlos Bernardos: Okay, let's move to the questions. Ayu? Try to be quick.

[00:46:08] Houyu: Hi. This is Houyu from FutureBee. Actually, I have quite a few questions. The first one is the recipient or the consumer of this notification is end host. So I need to confirm if this is actually for in the charter scope of this because we suppose consumer might be some network device. Because if you do this just inform the source host, it's become a transport layer problem. And, also, you as a second question is you send a faster CMP. The packet still continue to the receiver host. Receiver mice again process the ECN gain, which means you need to update the entire end to end RoQ e v two protocol to make it work. And in your presentation, you also mentioned the DCI connection for the long haul connection. That will further complicates the problem because DCI will involve much longer latency, which is not intended for the ROCCI to process the ECN because the ECN transport port protocol is fully has its own established to to process ECN to reduce the rate. It's a Hulk assumption. It's a very small very low latency within the DC fabric. If you consider the DCI, the problem is totally different.

[00:47:38] Faye: Thank you.

[00:47:39] Xiaomi: Let me quick I think that this is within the charter. If you see the slides in the first presentation and host as the consumer is within the charm charter. And your second question, you'll see that DCI is long distance, is long latency. But the as in a real deployment of the AI DC training, we need to have DCI involved. So that's the requirements. You can talk to operators. They will tell you that. Okay. That's it.

[00:48:18] Jeff: Just to ensure NVIDIA, two points. One, switch need not knowing anything about QPs. If you read IBTA spec, it clearly says that in CMP, you need to reverse source QP to destination QP. There are no remapping whatsoever.

[00:48:38] Xiaomi: So your question is about fast CMP. Right? It's not a question. It's a statement. For CMP, there's a as I as present the format of CMP, there's a field source copy within the packet CMP packet. So the switch if the switch need to send fast CMP directly, the switch need to know the source cookie. That's a must. Absolutely.

[00:49:06] Jeff: You parse it from the source packet and you reverse it into destination. It should be consumed by receiver to feed it into particular queue.

[00:49:14] Xiaomi: No. Receiver is not involved

[00:49:16] Carlos Bernardos: Okay. I see.

[00:49:17] Xiaomi: In CMP. Sorry.

[00:49:18] Jeff: You absolutely need to read a BTA spec. That's number one. Number two, reliable connections don't include specifically, don't include source. You need to figure out how to deal with it as well.

[00:49:31] Carlos Bernardos: Let's for the same thing, let's move that to the mailing list. Thank you. I think that very comments, but we need to move on. Thanks.

[00:49:39] Xiaomi: Thank you, Jeff.

[00:49:51] Jun Xinha: Okay. Hello. I'm Jun Xinha from China Unicom. Today, I'd like to present our work on use case and requirements for flow control collaboration across data center network and wide area network. This draft was presented at ITF one hundred twenty three and ITF one hundred twenty five in RTGWG. We highly appreciate the comments and contribution received from the Jeff, Injin, Dongjie, and Jinda. Based based on the main comments, we have made some updates. We added the appendix to provide a buffer requirement guidelines with formula and example. Second, revised section one introduction to further state the purpose of the draft and the cave alignment with the fine in conjunction network condition for DCI and this for DC and DCI. Third, as a working group established, so we moved this draft to choose a fine working group. For the background, with the rapid growth of the distributed AI training and influence across geographically separated DCs, lossless transmission demand expands from DC into one. PFC is widely adopted in DC, but as mentioned from the last presenter, it has the critical limitations were directly applied to one. So we need a fine grain flow control to to provide the fast and precise congestion control at a tenant level. The purpose of this draft is to describe the use case and requirement for interworking PFC in DC and with the FGFC in one. So we aim to enable the end to end lossless transmission, so the age node coordination and policy mapping of low control information between these two domains. This relies on the data plan fast congestion notification to meet the millisecond response. For attendance, we provide a buffer requirements for the network device in one to support the IGFC. The principle is that configure the back pressure threshold for for each tenant queue using the watermark. And when the occupied buffer exceeds the threshold, it will send a precise congestion notification upstream to stop or reduce the tenant traffic. We also provided a buffer formula and example at a 10 GPPS where over the varying one distance. This shows that a buffer depends on the wrong trip time, the congestion detection time, and the PIR of the tenant. Moreover, we also test several distributed AI training and inference tools in one, as shown in the below figures. This achieved over 95% computing efficiency compared with a centralized way. Both of these chose use the FGFC in y and PFC AIDC to guarantee the long RDMA transmission performance. And the gate eight and the h node at h node, the gateway made the requirements for the collaborative deployment successfully achieving the end to end flow control. We are now conducting further pilot verification with the health care and the finance industry clients in nine province of China. For next step, we are asking for more reviews and and we will keep alignment with the file and the related congestion control drafts. And we will further promote the network testing and validation to refine our draft thing. Thanks for your attention. Yeah. Question, please.

[00:54:14] Carlos Bernardos: Takumi, First, please.

[00:54:17] Takumi: Okay. I'm Takumi from NTT. So question on the buffer formula. It's two times RTD time times peak rate per tenant. At 500 kilometers, that's about 12 megabytes for 10 gig tenant. So a few 100 tenants means gigabytes of the output buffer at edge. That's expensive hardware. What's the assumed limit on distance and the tenant count? I think those number belong in the requirements.

[00:54:52] Jun Xinha: So so the question is so your question is what? Could you please sorry.

[00:55:00] Carlos Bernardos: I think it's more a a comment than a question. Right?

[00:55:06] Takumi: Question. So how how tenant your you you can

[00:55:15] Jun Xinha: How can I do this? You know, long distance? How how long

[00:55:22] Takumi: distance?

[00:55:23] Jun Xinha: How long is the distance, right

[00:55:25] Jeff: Yeah.

[00:55:26] Jun Xinha: For our trails? We have done the 300 kilometers, yes, from the two DC data centers. We have done this too.

[00:55:36] Takumi: Mhmm. Yeah. But some tenants

[00:55:40] Jeff: Sorry. Sorry.

[00:55:41] Jun Xinha: Sorry. We can talk offline for details.

[00:55:44] Carlos Bernardos: We don't have time for the last one. Thank you. Okay. Faye, I think you are remote, so I'll pass you the control.

[00:55:56] Faye: Yeah. Hello, everyone. I'm from. Today, my presentation is about the requirements for stability guarantees in packet spraying networks. Okay. Many packet spraying schemes have been proposed to mitigate loading balance in data centers. However, packet spraying brings challenges for stability guarantees. First, packet pass tracking becomes more difficult. In flow based networks, packets with the same five-tuple traverse the same network pass. By collecting anomalous packets and replaying their fab tubals, their network passes can be obtained for fall location. However, impact spring networks, the replaced packets with the same fab tuple may traverse different passes. Second, packet spraying networks require faster fall notification and convergence. Network failures have a more severe impact on packet spraying networks because packets in a single flow are spread across parallel passes. Therefore, packet spring networks require real time phone notification for rapid network re recovery. Third, packet spring networks need explicit loss signals. In flow based networks, packets from a single flow traverse the same network passes. Therefore, out of order packet is a clear loss indicator. However, packed spring inevitably causes auto order arrival. As a result, when packed loss occurs, the receiver must wait for a time out before triggering retransmission. This significantly degrades network performance. For efficient sorry. For efficient stability guarantees impacts bring networks. The first requirement is inbound anomaly detection and location. Firstly, packet loss can severely degrade the throughput of AI services and is also the most obvious indicator of network failures. Therefore, it's essential to precisely locate packet loss in real time. In flow based networks, operators can replace the lost packets to obtain their passes for fault location. However, impact spring networks, packs with the same fat tubal may randomly spread along parallel passes. This makes replay based packet tracking infective. Package screen networks require to detect and locate anomalies in band without replaying packets for pass tracking. In addition to loss detection and location, inbound performance measurement is also important. AI services are highly sensitive to network hotspots. In flow based networks, this phonetics can be located by tracking the packet passes with high latency. However, since the packet replay is inefficient in Spring networks, inbound performance measurements are also required to accurately identify and locate network performance bottlenecks in real time. The second requirement is fast fault notification and convergence. Impact imperfect load balancing networks, packaging a flow are randomly spread across different passes. As a result, the number of flows on each switch and link impact a spring network increases significantly compared to flow based networks. When a link or switch fails, a great number of flows are affected. Therefore, packet spring networks require faster failure detection and notification to achieve rapid network convergence. Furthermore, for flow for host based Spring schemes such as MRC, fast notification to source hosts can avoid Spring packets to 14 switches or links.

[00:59:36] Carlos Bernardos: Sorry to interrupt. We we have, like, twenty seconds so you can wrap up. Okay. Please. Yeah.

[00:59:45] Faye: The the third requirement is explicit pass loss loss notification. Sorry. Explicit packet loss signals. In flow based networks, packets within a flow traverse the same network pass. When auto auto packet arrives at the destination host, it can detect the loss event and notify the source snake for transmission. However, with random spring, all of auto arrival is inevitable, and the destiny device cannot identify by the out of auto arrival is caused by pack loss. So the explicit pack loss signals is very important. Furthermore, if the source NIC can be directly notified, retransmission efficiency can be further improved. That's all for my today's presentation. Thank you very much.

[01:00:29] Carlos Bernardos: Okay. Thank you. I I closed the queue because we don't have time. So thank you very much everybody for attending. I think this discussion should continue in The Middle East, and the purpose was to feed the gap analysis exercise that we have to do, and we will follow with the working group adoption call on the first item that we have. Thank you very much.

[01:00:49] Xiaomi: Thank you.