**Session Date/Time:** 23 Jul 2026 12:00 [00:00:05] **Adrian Farrel**: Alright. Third. [00:00:13] **Peng Liu**: K. Let's start our meeting. Good afternoon, everyone. Welcome to the CATS meeting in Vienna. We are Peng and Adrian as the co chairs. So first is a note where we are in the IETF process and policies. Please behave in a professional manner and extend the respect to others, same in the guidelines and the policy. And if you have any concerns, please contact the team. And if you are aware any of the IETF contributions in the IPR terms, please disclose the fact. And for details, you may find in this page. And the writing and audio video and the photographic codes is public. So if you have any other questions, please talk to the chairs and all ADs. So here's the meeting tips. Make sure to sign into the session via data tracker and use media echo to join the queue and show of your hands. And please keep audio and video off if not using the on-site version. And please help to take the minutes to make sure that your comments are recorded correctly. So here's the working group document status. As we have a new working group document of the OAM framework, and the CATS framework and the use case is in the RFC EDQ. And for the CATS metric definition, it's in the early review. So we we want that everyone could see the charter. Here, we're going to see applicability of existing tools and mechanisms, including that analysis of implementing the cast control and data plan, study potential new approach for the CATS control and data plan solution, and also study fault configuration, accounting, performance, security, requirements, mechanisms. And for the new OAM framework draft, it's matched charters now. So so we have some new milestones that was permitted maybe two or three years two or three days ago. And thanks to Jim. And you can see the milestones. We have data model document, and we have the applicability of the protocols and also for the OEM. K. So based on this, here's the agenda of this meeting. We have two working group document, presentations, and also have eight individual documents. And for the last one, please pay attention to the time. If you have the time, you can present it in the meeting. Otherwise, you please utilize of the mailing list. So any comments on the agenda? Okay. So we go to the first one. [00:03:56] **Adrian Farrel**: While Peng is sorting out the slides, I'm just going to put a question, a show of hands question about IETF-one hundred twenty seven, the next one. So this is a do you know your travel plans for IETF-one hundred twenty seven? Do you plan to travel? And it would be really helpful to us if you could just scribble that. Go on and find the next slides. [00:04:33] **Peng Liu**: So we got the first presentation. [00:04:58] **Kehan Wang**: I mean, the the box is still ongoing. [00:05:02] **Adrian Farrel**: Oh, I'm still [00:05:03] **Kehan Wang**: Thank you. Hello, everyone. I'm from China Mobile. So I will quickly introduce the progress on the CATS metric definition draft. So this is the the current version is a tense revision. So I think authors has been making progress in the from IETF one to one to five. Sorry. It's not working. Oh, it's okay. So we have a I mean, here we special thanks to Adrian for chair review of the document, And we also thanks to, like, Jack Jack Line for sec early review and the following for ops area early review and also, like, other review comments from the mailing list. These comments are really helpful in building this document. And, yeah, I think the document is approaching mature. And we have, like, four revisions since last ITF meeting, and there are some major updates. So the most important one is clarification on the definitions and usage of aggregation functions and the normalization functions. So I I recommend everyone who's interested interested in the work to read the latest version of the document. And, also, we have some clarification on the definition of the metric levels, especially on level one, and also, like, metric fields re refinement and also, like, improvement on security considerations. So regarding to the clarification on aggregation functions and normalization functions, it's very important because the these functions I mean, the definition of these functions, they are important to read each metric levels. And before the revisions, there are some, you know, ambiguous understanding about the these functions even within the coauthors. So we have many, many meetings, and then we have achieved some consensus here. So please see this simple example about the cons a confused concept. So we are assuming that you there are two types of level zero matrix, for example, CPU memory or CPU n, CPU frequency. They are chosen to derive a level one metric, for example, level one computing, and through some WittiSum. So and here, you know that CPU memory and CPU frequency, they have, like, units. And then we apply a weighted sum on these two level zero matrix, and then we generate a level level one computing metric. So this document argues that the above example belongs to aggregation, I will give you our explanation here because someone think this belongs to normalization. Here is the basic rules and examples for aggregation functions. So before this document, for example, RFC 5835, a metric combination framework defined in IBPM before. So it has defined spatial aggregation and temporal aggregation. So, basically, those two functions are applied on homogeneous inputs, and the output retains a physical unit of the metric. And this document further defines cross category aggregation, which means that the aggregation is applied on heterogeneous inputs, multiple input, and then generate one input, which is uniless. So I think this document specifically defined this. For example, some example functions are applying this aggregation function, like mean average and the mean of mean max and also weighted average. And normalization functions adjust orthogonal to aggregation functions. Here's the rule. So I think, basically, you just need to remember that normalization functions, you have a one input and then you've got one output. So when we convert a metric value here, whatever it has unit or not, it will be generated into a unitless and with the boundary score so it can be compared across heterogeneous resources. So some example functions can be like sigmoid functions or min max scaling, and then you can derive one input into a bounded score. And we can apply normalization to many levels. For example, you can normalize in a single level one metric to generate either a level one metric or a level two normalized metric. Or you can normalize the output of aggregating multiple level zero matrix and then generate a level one normalized matrix. So I think you better read these changes because they are important in this document. And then, correspondingly, I mean, some slight changes on the metric definition here, especially on the level one. Previously, its terminology is normalized matrix and categories, but now we change the name to be matrix combined in categories. And because the key distinction is yet is that level one matrix may retain physical unit if you apply spatial or temporal aggregation on level zero matrix, or the unitless, for example, if when you apply cross category aggregation or normalize value to generate a level one metric. And then we have added some implementation policies here. Like, we have added, like, score meanings and score semantic mapping here. And, also, we provide some recommendation on how you implement this normalization functions. And, also, like, we have recommend to you know, when you implement, you should have a measurement window, you know, for better policies. When we I mean, we have received some feedback from ops early review that this part can be, you know, put together into the operational section, so we will make a revision after this meeting. And, also, we have, like, updated version when we received some feedback from security area early review. So, basically, here in this document, we have, like, focused on five aspects of the security. It's a metric message or metric values. They are integrity, authenticity, controllability, freshness, and confidentiality. So in terms of integrity, we have, like, some consideration on cryptographic protection of the metric message, and also, like, you need to verify on the receiver side. And then authenticity is about, like, we have some, like, requirement on the metric publisher and the metric message authentication. And about controllability, we need some apply. We suggest implementers apply fine grained authorization policies for metric publications. And about freshness is also an important part and then make sure that the matrix is not staled. So we need some staleness handling about the matrix values. And then comes, like, confidentiality. So all the matrix cast matrix values should be, like, encrypted when they are either in transition or be stored. So for the next step, we will I mean, after this meeting, we will respond to ops directories review comment and then publish a revision soon. I mean, we will put together some operational, like, implementation policies and already some and other inputs from existing documents, I mean, which will be presented later into this operational consideration section because this document is standard track. And then we will start working group last calls. If you have any comments regarding this document, please feel free to listen this time. Thank you. [00:13:50] **Peng Liu**: Oh, is that okay. [00:13:53] **Adrian Farrel**: That's me in the queue. Hi, I'm Adrian. There are two other documents on the agenda today about metrics. Exactly. Yeah. And obviously, we we need to listen to those presentations. Are you thinking that there may be work that needs to be folded in here or kept separate? [00:14:16] **Kehan Wang**: Yeah. I mean, later, it will they will be presented, but I have, like, I have, like, before this meeting, reviewed this drafts. And and, also, I talked to some of the authors. I think their drafts may can provide some valuable input, especially on implementation and operational side. So I think this document is about to, like, add a specific input operational considerations section. So I think that those are drafts. Some of ideas from those drafts can be a good input to that part. Thank you. I'll comment later. Thank you. [00:14:59] **Peng Liu**: Okay. Thank you for the draft for the authors of the two other metric draft. Please keep in mind when you're presenting a work. So next. [00:15:28] **Juan Deng**: Hello, everyone. I'm from CT. I will pre presenter the CATS OAM framework behalf of on behalf of their coauthors. This is the updates from last versions. We have presented at the last meeting and the last meeting, and thank you for comments and suggestions from these people. And your feedback are very helpful for us. So we submitted some versions from version five to six, seven, to zero zero, and zero one. And we made much revisions. For example, we removed the terminology such as the SI-OAM and SVC-OAM, and we added the new devilutions, push mode, and the pull mode from the comments the comments from the. And we also create a finder clarify the modification of problem statements, clarify the gaps, and the the rig the updated the related reference. And we also revised the the layering model and the components to align with their align with the names between the layer model and the components. For example, the instance OAM and the service OAM and remove the SIOM and the SVC OM. And we also clarify the end to end trace mode from the comments is from the. So we also revised the sum add added some CATS OAM requirements, refined their requirements across operations, configuration parameters. We also added some the the edge section of the deploy deployment considerations. We also added some security considerations for OM specific. So let's recap all of this CATS OAM framework. First, about the it's about the motivation and the problem statement. Why we lead our new CATS OAM framework? We think that maybe there will be maybe some problems. For example, the traffic may be selected from the to to their to to be collected to their subs instance, and we need to consider their computing metric, and we need to low locate the the end to end of force, including the subs instance. And we we need to very verify their network and the computing metrics. So in our in one word that the CATS OAM, we has their key requirements to ensure the best inch their instance children by their cast for to actively guarantee the reachable and performing. So their of the traditional OEM monitoring just stop at the egress router, and we need to extend it to the service instance. So this is the motivation for the CAS OEM. And so this is the relationship and their place of their CATS system. We have CATS framework to define some to define their architecture and some key components, and we have use case and requirements for the CATS service. We also have the metric. And so how to monitor and ensure their the end end the the end end to end service and how to monitor the metrics. So this is the relationship with the other documents. So this document leverage the existing ITF OEM standards. For example, the RFC seven two seven six, the OEM tours, and the the draft of operation operational RFC five seven zero six base and RFC six two nine one and the and the other. We will mention the other existing measurement tools, such as the PFD and TWAMP, stamp, and other other measurements tools. But we also clarifying that this document has has the as best defined low protocol extensions in their associated tension or lot of their scope of this document. So first, we propose CATS OAM layer remodel. It it is four layer models. For example, the service service OM, their instance OAM, and the path OAM, link OM. And the path OAM and the link OM, we can use the existing tools and their existing functions. And there there we added some two new instance, no mode OM and service OAM, the two layers and their corresponding components and the functions. For from their comments from from Charles, we added their service OM for including end to end mode and the trace mode. So to specify this their component the functions of their two layers, we define their functional components with the same names. For example, the instance OAM, we need to provide the status monitoring, metric monitoring, and the reporting modes, such as the pull push mode and the pull mode. And, therefore, surface OEM component, we need to provide the policy policy verification and joint performance measurements in the multi domain fault isolation. So based on the OEM lay layering model and the corresponding components, we also propose some new requirements for the operation, administration, and maintenance. This some for example, for operations, we need to provide them their multidimensional dimensional status monitoring. We need to provide service ID and location, their telemetry integration with the CSMA and the computing metric procedures handling. And for the administrate administration, we need to provide a policy policy based steering configuration and the the differentiated monitoring and the the service and the metric integrity. For the last one, the maintenance, we need to provide a joint fault of perform sorry. And the the hierarchical choice traceability and the forwarding play the for pro constitution to check. And we also add our section for deployment consideration. We these are some examples for the subsOM and the instance o OEM at their their from their ingress CAS forward to their target SI. And, also, we list some example for configuration parameters for subsOM and the instanceOM. And we also add their security considerations for protection of OM metric reporting channel. For example, the OM metric reports often often authenticated and integrated protected under the reporting channel provides the con confidentiality of the service information. OEM control message should be protected against the the the attack. And therefore, OEM specific consideration, the I o IOM master not leak their client and then if they call information and the follow secretion. The security guidance will be as per the r c nine one nine nine seven. The data will collect all in the metrics less read for the telemetry. So this is where the the updates of their CATS OAM framework and further review and the feedback are very welcome. And the next step, we will align with the CAS framework and their CAS metrics doc document. Thanks. [00:25:34] **Peng Liu**: Thank you, Luis. [00:25:36] **Luis M. Contreras**: Thanks, Juan. One one question here. Because if I understand well, you you are considering that the service instance will have the the possibility of responding, for instance, that e one flow or whatever. So because the service instance is running in a virtualization environment, I was wondering if it will be simpler to solve this just basically installing a prop in that virtualization system collocated with the service systems that could manage the OIM. Because otherwise, we are imposing requirements to the service systems that is something out of our control, let's say. So my point is, wouldn't it be simpler to use just props and and that's all? [00:26:15] **Juan Deng**: Yeah. Yeah. I agree with you that we that the the OM monitoring just in the the report, just monitor the the the metric report, not monitoring their instance inside the intern instance. [00:26:33] **Luis M. Contreras**: Instance. [00:26:33] **Kehan Wang**: Yeah. Yep. Yep. Okay. [00:26:39] **Sheng Li**: From Huawei. So first of all, thank you for your work. I have reviewed your document, and sorry for my late to reply your email yet. So, however, it seems to me that the draft might need more content. Right now, I don't see so much new content specifically related to, you know, cat. So I hope in the future, you might add more tags to the OAM mechanism, yep, in the draft. Let's make it a useful draft. Thank you. [00:27:13] **Juan Deng**: Yeah. Thank you. We have made a new revision to your comments, and please feedback more give me more feedback for the new revision. And what you what you just saying is that we lead to mention more OMOM specific tours or existing mechanisms. We will do more clarification next version. Thank you. [00:27:49] **Peng Liu**: Okay. Thank you. There's also comments in the chat box from Boris and. [00:27:56] **Adrian Farrel**: Is it? [00:27:57] **Peng Liu**: Oh, okay. Okay. Boris, [00:28:01] **Boris Khasanov**: Thank you, Boris. I'm just wondering in regards to tools, why you did not consider stamp additionally to to AMP or instead of to AMP also? Because it looks more, like, more advanced protocol for that purpose. [00:28:21] **Juan Deng**: Yeah. Yeah. We just I think we just provide some consideration for the measurement because is a the existing tools for measurement for their metric measurements, and then maybe it will be extended for their computing metric and for their surface instance. [00:28:45] **Adrian Farrel**: Thank you. Okay. So the chairs rather rationally asked the AD to add a milestone for an OAM framework. Does is there anybody out there that believes we should not be looking at OAM for cats? Nobody is brave enough to go to the microphone. I think there's work still to be done on this document, but it it it's moving towards being the right thing for our foundation. So keep working on it, and we'll get around to adopting it. [00:29:23] **Juan Deng**: Okay. Thank you. [00:29:33] **Peng Liu**: So next is protocol applicability. [00:29:53] **Linda Dunbar**: Okay. So I'm presenting this on behalf of my coauthors and on the protocol applicability for CATS. So the goal is to analyze the protocols for CATS and where the extensions needed and where not to use them and if there's any limitation. So as you see that cats itself need a suite of protocol. It's just not just for one thing. Right? So there's a matrix collecting the matrix and for distributing those matrix. And there's collecting the matrix. And there's also steering of the traffic based on those matrix. So there are quite a few different categories of protocols needed for for cats. And the goal is really this particular draft is to look at all the available tools, protocols available, and see which one is suitable and which one is not, or if there's any limitation when use them for cats. So start with BGP. So the BGP is for, like, really useful because it has the policy control. They can control where to distribute them, which ingress nodes can receive them, and how do you distribute them. So they are useful for limited domain. You really don't want distribute all those metrics to a different domain. Those metrics are not supposed to send to them. So BGP has really good policy control to constrain them, and there have many available mechanism to do this. And and but BGP is not really suitable for distributing, like, l zero matrix defined by CATS because there are so many of them. And, also, they fluctuate. They changes all the time. And so BGP is now really suitable for distributing those matrix which changes very rapidly. So it's appropriate for l one, l two type of matrix. Best is this aggregated value will be the best. So here, we have in IDR. This draft has been the in the working group draft for a long, long time, and, hopefully, it will start working group last call. So over there, in idea, I would define this pass attribute. We call it the metadata pass attribute. And this pass attribute basically attach with the the routes and to indicate all the associated metrics for associated with those routes. Specifically, there were, like, five categories of sub TLVs being defined. And there's preference, site preference, which is aggregated value. And there's site availability. That's physical availability. I think cats has a similar draft. It's aggregated. Basically, you have many, many services hosted in one part or one location. If that location lost fiber fiber cut or lost power or something happened to that particular location, all the services, l force I mean, IPv four, IPv six, the MPOS, all those services are impacted. So in the IDR draft, we have a simple mechanism to indicate those services that basically has a metadata associate all those services with common physical characteristics. So when something happen, the the nodes, the router doesn't have to send individual route withdrawal for each of them, and they can send once combined message. There's okay. That's the physical availability. There's also there's a service capability, like how this resource the how the resource in this particular site can accommodate those kind of services. You can read through them. They have five sub tier ways defined. Next one is so so the BGP I just forgot to reach the conclusion there. BGP is suited for limited domain and is for values matrix which doesn't change very rapidly. Next one is the IGP and BGP LS. So IGP, this particular to distribute is mainly for distribution, distributing the the matrix. There was a draft discussion in the l s r LSR working group. And the conclusion from there is saying that, basically, because IGP is processed is a flooding mechanism. It's processed by every nodes along the way. And for cat matrix, like, from the source of the matrix to the ingress node, which need to use the matrix to steer the traffic, lots of the intermediate nodes may not even support those matrix or processing those matrix. They don't know what to do with them. And, also, that because flooding will flood to all the nodes. So it's not really suitable for distributing the matrix, especially the matrix, which is, like, l zero, l one. But it's suitable if you have aggregated value to indicate aggregated value on how suitable this particular sites for this particular services. And then there's BGP-LS for the topology. This can be useful for the past selection. Next one in the draft, they discussed the PCE. I think this particular section, we need to send to the PCE working group to get some feedback. So far, we haven't got much feedback yet. It's really for the past computation, How to select the path and be able to utilize those metrics to be able to use a centralized location to compute the past sent to the egress node. Here, we there's a section in the document talking about management and management plan applicability. So there's a young module for configuration and there's a control plan, which we just covered with those being analysis in different sections. And there's a data plan. So general analysis conclusion is those policies should not be encoded into the data plan. It will be too much to go through the network. Any questions? [00:38:07] **Peng Liu**: Please. [00:38:11] **Aihua Liu**: Hi, Lina. Thank you for your presentation. So, we also have a work for a user BGP-LS to collect the computing resource. And so we could have, do the analyst and consideration for, the more detail and to contribute to this applicability draft. And another question we think in the cast system, this is more capacity such as centralized or the hybrid. So you list you just list the protocol one by one. So how to consider the more multiple protocol, how to collaborate to achieve something? We suggest to maybe aid a section to analyze the how to use multiple protocols to achieve one as it's a very complex CAN system. [00:39:02] **Linda Dunbar**: Oh, that's interesting aspect. So you're saying that we have BGP enabled. We have IGP enabled. And then how does those two protocols collaborate with each other? Yes. In terms of distributing the matrix? Okay. We'll come input. Maybe you can suggest something to for the document. That'll be great. [00:39:21] **Aihua Liu**: Okay. Thank you. [00:39:22] **Peng Liu**: Yes. But this draft first is to analysis work about the protocols, but not to recommend any protocols or any solutions in this draft. So I think maybe some analysis work to for the questions of is valuable. But please don't to to any specific solutions in the draft. Yeah. So Sean [00:39:52] **Sheng Li**: Lee from Huawei. So thank you for your work. I think it's in I think it's, you know, important work, but we might need to, I'll say, ask more experts from BGP and IGP and BCP experts to to contribute to the work, to consult them. What do you think about this? Like, personally, I don't really think PCP is a good choice for the solution, but that's my own opinion. So, yeah, we might need more discussion between different among different working groups. Yep. Thank you. [00:40:28] **Linda Dunbar**: Thank you. That's a good suggestion. [00:40:31] **Adrian Farrel**: Yeah. And and do remember this is, at the moment, an individual draft. As it assuming we progress something into the working group, then, yeah, we we have a big responsibility as cats to talk to all the other working groups and see whether we are misrepresenting their protocols. [00:40:51] **Linda Dunbar**: That's a very good point. Thank you. [00:40:53] **Adrian Farrel**: Maurice. Thank you. More questions, Linda. Boris [00:41:00] **Boris Khasanov**: Khosanna from MWS. Linda, since you mentioned young models, why don't consider, for example, sort of telemetry, either or which is developed in a. [00:41:18] **Linda Dunbar**: That's a good point. Maybe maybe we need to add that analysis into the document. Thank you for the suggestion. [00:41:24] **Boris Khasanov**: Thank you. [00:41:27] **Jim Guichard**: Hi. Jim. I'm talking as an AD. I I thank you for the work. My concern is that we I don't wanna see the working group jumping the gun. What I wanna see is a document that outlines what the requirements from a protocol perspective are for cats. [00:41:51] **Linda Dunbar**: Okay. [00:41:52] **Jim Guichard**: And not make any pre assumptions on what those protocols should be. [00:41:57] **Linda Dunbar**: Okay. Thank you so much for those [00:41:58] **Jim Guichard**: details. Thing that I wanna see or rather I don't wanna see is a proliferation of documents in other working groups defining things on behalf of cats. If I see that, I will certainly jump in. [00:42:14] **Jeff Haas**: Okay. So [00:42:15] **Jim Guichard**: just a just a that's more of a warning for the working group because I'm already starting to see documents turn up in IDR. I do not wanna see too many of those until the working cats is done. So if you can focus on, okay, what are the requirements based on the framework that will become an RSCs in the RSC edit edit queue now? What are the requirements? And and by the way, the metric document's [00:42:40] **Luis M. Contreras**: not done yet. Mhmm. [00:42:42] **Jim Guichard**: So we need to know what what are these metrics. [00:42:44] **Luis M. Contreras**: Okay. [00:42:45] **Jim Guichard**: And so focus on what are the requirements for on the protocol rather than what protocol might we be able to use. [00:42:53] **Luis M. Contreras**: That that [00:42:54] **Jim Guichard**: would be my my preference. [00:42:56] **Adrian Farrel**: Thank you. [00:42:57] **Linda Dunbar**: You. So that's on the meeting notes. Right? So I will be able to if I forget, I'll be able to retrieve Did someone say IDR? [00:43:05] **Jeff Haas**: Someone did say IDR. Jeff speaking with the IDR chair hat. No. At least one of three. So first of all, Linda, thank you. You have managed to capture most of the protocol behaviors, you know, from your five g experience well in the document. I think that this will be helpful for the chairs and the working group. I was also standing up here, and Jim has made the majority of the points that I would come up here to make, which is that even having captured it, you know, the five g documents a specific use case for a cats like behavior, you know, that's moved forward because of the order of how things were chartered. Emphasizing what Jim's saying towards the chairs, what I would like to see, at least myself, for know, IVR is that we focus on what type of behaviors you're trying to accomplish, get that refined within the context of this working group. [00:44:00] **Luis M. Contreras**: Mhmm. [00:44:01] **Jeff Haas**: And then when you're finally to the point that you have a sense about what information do you want to put into the protocol, maybe invite people to come here, know, to do the work initially. And then once it's matured a little bit, take it to IDR. This will avoid some of the pattern we've seen with other technologies like Sav and other things where people start shopping the protocol before the working group is actually converged on its requirements. [00:44:25] **Linda Dunbar**: Thank you. Thank you so much. That's it. [00:44:31] **Adrian Farrel**: Thanks, Linda. [00:44:32] **Linda Dunbar**: Thank you. Thank you so much. [00:44:43] **Peng Liu**: So next, data model. [00:44:56] **Juan Deng**: Hello, everyone. I'm from CT. I will talk about the date date model for CATS on behalf of the co authors. To answer the questions from Charles, first one is why CATS needs a data model? From my understanding, we have a CAS framework. And for from this document, it requires the the CAS management plane. And so it is responsible for the monitoring, configuring, and maintaining their CAS network deficits. So and the CAS data is requirement to to be maintained in the management plane and configured to their control plane and data play and their CSMA. So this is I think this is motivation for that model for CATS. And there are also we to answer this question, a young data model is required to structure, validate validate, and what to make this date across the CAS components. So we just asked the presentation in a presentation that the cast data model request the functionalities such as the configuration management and traffic steering policy configuration and their service metric management and their their refund reporting. So we will provide their young data model for this requirements and their funk functionalities. And the the second question is how their data model will be used. The data model will be used for CAS information, maintaining, and their configuration in interfaces among their management plane, control plane, CSMA, CNMA, and the CAS forwarders. So their document defined some information base. For example, their CACI CIB CAS computing information base to responsible to for maintaining their network computing information and provide some base data for CSMA. And the CNIB is responsible for maintaining their CATS network information provides basic data for the CNMA. And we also proposed two interfaces, the CATS SBI interface to it it is used to report computing metric information from, for example, from their CAS forwarders to their control plane, or it also can be used to send their their pass and the service policy information as of or surface information from their controller blade to the CAS forwarders. And the the CSMA API is used to report their report their their surface metric and weight information from their CSMA to their control plane or their CAS forwarders. So this this architecture is not aligned with the framework. So we have presented for many times, and thank you for the comments from people. And they are very useful. And we also made some revision from the fourth version to the seventh version. We have completed the the structural reorganization for clarification, and we also clarified the young data model scope. And we modified the young tree to align with the RFC eight three four zero and revised the young model model model based on the data model modification and modify the terminology. And there we supplement their security consideration and their the IANA consideration section. And this oh, it doesn't work. And this is the first update for the data model. We changed the traffic classifier container as the traffic class classifier list. This is to address the comments. And during their traffic classifier list, their protocol type is changed from their their 16 to to to the eight. And we also asked the read only statistical counters. And this is the tree of the track cast traffic classifier tree. And the second update is that we are in the metric list, the metric leaf is renamed to the metric value, and we also ask the read only the read only state decode counters. And there this next one is the in the load notify notification, the entry limit reached is related to the metric limit reached. The young tree syntax has been modified to align with the section 2.6 of the r C8340. So we have made many revisions to address the comments, and we have re re resolved them the order comments received. And we think the workforce thing is a study for request working group adoption. Thanks. [00:51:08] **Peng Liu**: Okay. Thank you. No way in the queue. [00:51:11] **Adrian Farrel**: Can you go back to Slide number two? [00:51:18] **Juan Deng**: Which way? [00:51:19] **Adrian Farrel**: Which way? Number two. [00:51:24] **OTN Presenter**: Two. Okay. [00:51:27] **Juan Deng**: This one, how the data model will be used. [00:51:29] **Luis M. Contreras**: One more. [00:51:32] **Adrian Farrel**: That's it. [00:51:33] **Juan Deng**: Yeah. [00:51:33] **Adrian Farrel**: So question for the room. These first two bullets here lay out who's responsible for the data, the the sort of management data, and then boldly asserts a young data model is required. Does anybody want to speak against those assertions? Raise your hand if you're still awake. Yeah. That's not so many. Okay. So question for you. Do you think well, we have a yang doctors team to help us with yang work. Do you think that this is now stable enough that we are ready to use them to help us make sure we're in the right direction? Or do you want to work on it more before we use them? [00:52:37] **Juan Deng**: I think this is stable for for our use. And we have made many revisions, and we align with the framework and the metrics. So we Okay. [00:52:49] **Adrian Farrel**: Yeah. Okay. Then the chairs will ask the young doctors to have a look. [00:52:54] **Juan Deng**: Okay. Thank you. [00:53:00] **Peng Liu**: Okay. Thank you. [00:53:12] **Jeff Haas**: Sorry. While we're waiting to change the speakers, for the Yang doctor sorry. Jeff Oz again. For the Yang doctor's review, a fair warning, their primary job is to help with structural and also syntax related reviews. They're not very helpful for semantic reviews. It's incredibly important that other people take a look at that aspect. [00:53:36] **Peng Liu**: Thank you. [00:53:42] **Aihua Liu**: Okay. Hello, everyone. I'm Mo Young Han from China Telecom. And today, I'm going to present this work, distribution of service metadata in BGPIS. So firstly, I want to see this work is from, there are some overlap with cats. So for here, I want to talk the requirements of why we use BGP-LS to collect the service metadata and to distribute the service metadata. Okay. First, we have a quick recap. For the ATF-one eighteen, we present the hybrid solutions for the computing awareness theory. And for the 01/2020, we discuss the protocol applicability and share the China Unicom deployment experience. After the meeting, the hybrid architecture was referred as the one of the cats' active options. And for the item one twenty four, we explain why we use PDP and DUS are useful for the cats and introduce the protocol extension details. And the draft was updated according to the discussion, including the alignment with the cast metric work. And in this year, in the interim, we present the requirement to protocol extension and the law in ADR. But, yeah, as a design, so and ADR give us a comment when should to firstly to clarify the protocol requirements and the availabilities in CAS first. So I'm here. So next, I will to show the our proposed solutions to collect the computing aware information. So this is a proposed centralized collection model. So there are three steps of this workbook flow. First one is the collections. For here, it's a centralized model. We use BGP-LS to carry the service metadata from edge sites reported by the CSME and underlying topology northbound to the CPS. And then we do the processing and to the CPS, the path selectors, which may be co located with the PC or the SDN controller, combines the network and computing aware information to select the best set and the path. The third one is distribution. The selected policy or the path can is stored through the separate mechanism such that the PWL flow spec or the SIR policy depending on the deployment. So this is our proposed one solutions. And based above solution, we have to do the deployment and the verification in the real scenarios. So this based on the above scheme, we have to for to build the current system for the remote driving scenarios in Shuang, Habit China. And this is for the detail about the experiment results, and you can find in our previous presentation and draft. Zen is a BGP-LS format data collection. I think now we talk the detail about this since it's too early. So I just to to clarify why we use the BGP-LS. So for the reason, I think is that BGP-LS already provide this northbound mechanism to collections network topology information. So this chapter to extensions mechanism to include in some service metadata needed by the computing aware of past computer. And another detail that I think is from the ADS work. Don't talk to them. And I think this is also the requirements we should align with the CAST discussion. And we have answered two questions and some related work. First one is why we use BGP and BPS. We think both are useful since for CAS, we have the centralized model, distribution model, and hyper run model. So for the BGP based distribution, service routers and a fast metadata delivery directed to the ingress routers. They support the distributing traffic steering based on the local policy and network cost and service status. And on another hand, the BGP-LS is based on the collection centralized collection. So the service router and the metadata can be report northbound to the CPS, and this support the centralized path computing and the global optimization. So these two mechanisms are complete. So BGP enables distribution decision to the best provider as a data channel for for the centralized decisions. And the yeah, I want to talk about the main comments for the rate of the change of the end of our assessment information. So this draft follows guidance to the idea of matching definition. So for the high frequency raw data, we should not to address the directly use of BPS. So we can carry the L2 or L3 metadata to avoid this other test. Are also lists some related work, I think, to we have some different reasons for the first one is from the in desk work in the group. They define the metadata model, and in our work, it's define how to report it to use the BGPRS. And second is a cat's availability draft has present early. So our maybe to mainly focus on the detail about the and we can collaborate to to do the further work to contribution. And for another work, we think is mainly focused on the distribution scheme, and our work mainly focused on the centralized decision in cats. Okay. The next step, we should also continue along with the draft of the cats' architecture metric and the data model. We can do, firstly, to show the requirements of why we use the BPS. And we also to do the further study about the rate of change and some security in our follow-up work. And we also need to coordinate with IDR's work and to according to the to update our draft. Okay. This is my presentation. Thank you. [01:01:42] **Peng Liu**: Thank you, sir. Please. [01:01:54] **Susan Hares**: This is a plea for the working group based on your comments. So please, this is a request for the working group as it relates to IDR drafts. So when we do an IDR draft based on another working group, it may not be really clear but we come back to this working group and say, is this draft necessary for the framework? Is this draft something that's defined and has a good piece? We often ask, you know, there's an appendix that says how it's working that way. So as you go through, if you go back to your previous slide, the one just before, you have a number of IDR drafts there. You have the five gs, which is a BGP attribute, and then you have other BGP features, BGP LS. So when they come back to IDR, they need to come with something from this working group saying, gee, we think as a whole it's a really good thing. So I'm asking you and the chairs to work toward that so that when I get these drafts and you're asking to publish them, I have the information. And perhaps if you're really kind, you put that in appendix that can be removed. Does that make sense? That's my request. [01:03:36] **Adrian Farrel**: Yeah. And to echo what Jim said earlier, if people are coming to protocol working groups saying, we need this for cats, They are premature. [01:03:53] **Susan Hares**: Gee, that would be good to know because then we will review them at that level. [01:04:01] **Jeff Haas**: Jeff, was building on what Sue has said. Using your slide right here, five g is a similar technology that is exposing level one type metrics. You know, that was early. For the next set of work, that very large list of drafts. One of my requests to this working group would be to analyze the different drafts to see whether there is common CATS level one and level two metrics that can be exposed selectively through either BGP, which will be per destination, or BGP LS, which will be no link state, and see if you can decompose this into common cats use cases. If you can successfully do that, you'll have a compact piece of work to take to IDR. [01:04:55] **Aihua Liu**: Okay. Thank you. Think for to talk more the requirements and for the why we use BGPIs. So for the detail about our experiment and deployment, we can to discuss at the mail list with IDR scope, I think. So and for yeah, I think today, to to talk how to extension and the the detail according to the source comment, I think it's too early. So some cats work is not to get the constants, and they may be too need time to give the more clear guidance for how to use the BGP or their give us the requirements, and then we can to get the detailed guidance and to design the protocols, I think. Okay. Thank you. [01:06:00] **Peng Liu**: Okay. Thank you. So [01:06:14] **Longfei Xia**: Hello, everyone. [01:06:20] **Lina Ji**: I'm Inna Dai. Today, I will present our new draft computing service metric definition and operation under CAT. This draft of meaning focus on how to mold the service side information or service oriented way so that the CAT's control plan can use it directly for traffic steering. Oh, our motivation comes from a practical crash issue in CAS metric dissemination. The CAS need steering traffic based on both computing condition and network condition. Network metrics such as bandwidth or delay are relatively straightforward for path selection. However, the computing metrics are much more heterogeneous and dynamic. The service side may have CPU, GPU, MPU, and storage. At the same time, the local load, healthy status, and reachability may change frequently during operation. If all those raw metric advice advertised directly, the CPS may receive a lot of data but still don't have a clear answer to the real steering question, which service instance should be selected for this request. Actually, we think the CPS need the service level information that is meaningful for selection, stable enough for operation, and directly usable by the control plan. We summarize this issue into three operational gaps. The first one is the implementation gap, how to normalize The computing results are heterogenous, and it's very difficult to map them in one sphere and the unified score. And the second is information loss gap. Even if we have a normalized score, it may loss service level meaning. For example, a a score of seven doesn't tell the CPS whether it can serve serve as more clients or it can satisfy a strict processing time requirement. The third is routing mechanism gap. Can SAP ask me steering parameters, not hardware specific results details? So our draft is to bridge bridge the gap between metric description and operation theory. Before we introduce our proposed metrics, I want to clarify how our draft differs from the existing CAS metric definition work. The metric definition draft defines the level one coverage metrics and level two metric for both communication and computing information. But our draft takes a service oriented view using CSID and CSID as indices. It expose service metrics such as gas and computing time. In short, the metric definition provides a general metric framework where our draft provide service specific operational metrics. Based on this difference, our key idea is service oriented abstraction. The service status still the service status still use the information, but they remain local. This also drives the question we received on the mailing list. How are the service oriented metrics derived from the basic metrics? Our answer is the service side derive them locally based on their locator results, runtime status, service requirements, and local and local policy. In this way, raw hardware, still details, they all remain local. The CPS receives the information that is directly usable by the control plan. This table shows the proposed metrics. CSID identifies the request to service. CSID identifies and locates the service context to the instance. The two key metrics are guest and the computing time. Guest represent available concurrent capacity for specific service. Computing time represents the expected processing time for one request. Together with the network delay, the CPS can evaluate the total service time. Operational metrics such as cost rep reputation, security label, and the capability can support additional deploy deployment and policy requirement. Derivation algorithm is local, and our drafts mainly focus on the metric semantics and operational use. This slide shows how the CAPS combining computing metrics with net network metrics for service instance selection, computing service table. It's built from the, say, SMA reports. The network service table opt ins existing network metrics from, for example, TDB in as in controller. A mapping between the service contacts instance with the with this egress helps CPS allocate the each candidate with their network pass. And when request arrive, the CPS first find the candidate on from the computing service table, then obtain the network metrics for each path network path, and finally, combining computing time and network delay to select the final CSI CI ID with the shortest service time. Our drafts reuse the existing network metrics instead instead of redefining them. So what I think service oriented metrics are useful. First, they are operational actionable. CPS can directly use gas and computing time for service selection instead of integrate from hardware's details. Second, they are hardware independent. Raw CPU, GPU storage, default details remain local and to the service sites. The CPS only receive the service level information. Third, they are capable with exact CAS metrics. Our draft is to complement for operation the the normalized metrics from for operational for replacing them. So, finally, they support scalable reporting. This also drives a question which we saved. Since the gas may change dynamically, do we need to report every change? Our net our answer is no. A small permission change can be handled by the service side locally. Well, the significant change can trigger the threshold based updates, and periodic heartbeats can prevent state. So to summarize, our next steps is to refine guess and the computing time definition, clarify threshold trigger updates and hybrid synchronization, and add more CAS deployment examples. Thank you very much. Feedback's are welcome. [01:14:31] **Kehan Wang**: Thank you for of this draft. I think, as I mentioned earlier, in cast matrix definition draft, I think it is a good input on especially on operational considerations, provide some you know, you have some implementation details. You can go back to the slide on the two metrics that you define here, especially on gas and I mean, sorry, the gas and the computing time. So when I it'll map that to the metric drop. Like, for example, the gas, can see, like, a level one service metric. Or computing time, you can see it has a level zero metric, and you apply some compostable policy on how to implement those metrics and make decisions. I think it's good. It can be an input input to the catch CAS metrics draft, but I I I don't recommend this document to be go separately because I think it's just one part of the CAS metrics draft. And, you know, CAS, we don't I mean, we will we're not define specific implementation. We just provide some guidance on how to use this metric. So that's my comment. Thank you. [01:15:41] **Lina Ji**: Thank you for your opinion, but we think this draft need to be separate because the I have told that the metric definition is more like a general metric framework, but our draft is service specific. And guess is metric that is directly for use, and we don't think it like the level one metrics. Thank you. [01:16:16] **Peng Liu**: So please. This [01:16:26] **Susan Hares**: is a set of informational comments and you may just want to point me to someplace in another CATS document. But it would help me to know if you're expecting that computing time is synchronized across a network. Are you expecting that the the machines are going to have the same sense of computing time, or are you simply looking for a delta within a network. [01:17:06] **Lina Ji**: We think that the service sites to measure the their computing time and report their operational computing time from us from like, if through their mechanism, they will measure the computing time on their own. [01:17:30] **Susan Hares**: Okay. I'm gonna repeat that back. And I appreciate your help because it helps me to review requests for IDR flow spec and IDR other things. You just told me that computing time is a delta on the machine that it's targeted for. Did I understand that correctly? [01:17:59] **Adrian Farrel**: That's that's what I heard as well. So this is an an elapsed amount of time on a device. [01:18:07] **Susan Hares**: Okay. [01:18:07] **Adrian Farrel**: So timestamp synchronization, not interesting. That's what I needed to know. Real understanding of how time progress progresses is sort of assumed, but devices may lie. [01:18:22] **Susan Hares**: Okay. Because I've had some that come in saying, I wanna schedule a specific time and other time ones that come in informing me about this delta, how much time it takes. Thank you for your answer. [01:18:37] **Lina Ji**: Thank you. [01:18:40] **Adrian Farrel**: Thank you. [01:18:47] **Peng Liu**: So that's yeah. [01:18:49] **Lina Ji**: Now I will move to the second related draft. It's about the public service platform for cats. The previous one is about the service metric. This one looks at the platform that maintain the reference context for those metrics. Let me quickly recap why we need a public service platform. The SAS ID in CAS identify the service, but identifier itself alone is not enough. It doesn't tell the client how to use the service and what data should be the input and what capability the service provide. And so it it doesn't tell the service side how to deploy the service. The service side may need the code location, recom computing or storage requirements, reference context. So our idea is to introduce platform as a common catalog. It makes case CSID relate to the contact to either to discover, and different anticipate can interpret it in a consistent way. [01:20:14] **Adrian Farrel**: Sorry. [01:20:17] **Kehan Wang**: It didn't work. Oh, [01:20:22] **Longfei Xia**: in [01:20:23] **Lina Ji**: this slide shows the where the public service platform sits in. The platform interact with three types of users. Client usage to obtain CSID and service context. Service data usage to select and deploy service. Service provider publish service entries to the platform. In cats, the platform act as a catalog and publication point. It provides service context before traffic during begins. It doesn't perform service instance selection or and is not a part of data plan forwarding pass. Services test selection and forwarding still follow CAS mechanism. This draft is mainly about the role of the platform and the information it maintains rather than a specific complement. Depending on the deployment, the platform can be realized using such as the AI in as based architecture, search engineering, AI assistant agent, or other architecture. This slide shows the information model and the platform maintains service interest like the example shown in this table. We summarize these fields into three grow three categories. The first one is service description. This field meaning help the help the clients to understand what the service is and help them to construct their requirements. The second is deployment context. This failed helpful for the service sites that want to host the service. The third is reference context and metadata. Reference computing time and reference guest to catalog level significant values. They give the service data start point before the deployment and measurement. After the deployment, the service desk still need to report their operational computing time and operational gas. After the abstraction and information model, I will introduce the logical workflow. First, in service publication, provider publish service interest to the platform. The platform assign or generate a c CSID. Second, in service side deployment, the service side queries a platform for deployment for deployment context and then deploy service instance locally and report operational metrics after deployment to use it say SMA. And third, in client service resolution, the client should submit their service description or requirement to the platform, then platform resolve it into one or more candidate CSIDs and ways related to service context. Here, the resolution doesn't mean to find the final service context instance. Clients use the returned service context to to construct their requirement, and then the existing CAS mechanism use it to find the service context instance. This workflow is to expand how the information is used and they do not define a mandatory protocol. To summarize, our next step is to refine the information field based on the feedback, clarify the relationship with service oriented metrics, and collect more use case for service publication and resolution. Feedbacks are welcome. Thank you very much. [01:24:42] **Peng Liu**: Thank you. Yun Zi, please. [01:24:46] **Tsinghua University Representative**: Hello. I'm Yun Zi from Tsinghua University. I have a brief question about that. I see that there's also a down buff, and that discusses the discovery of agents and their workloads and some things. What you think is the relationship of your work and down, and which aspects are kind of speak, and which aspects that can using that can coordinate with what? With DOM? Yeah. Thank you. [01:25:14] **Lina Ji**: Thank you. We think that DOM is more about the agent discovering. Well, our platform is about the interest that can be deployment on the cats, and we can give the we can most suitable we are most suitable for CAS framework. Thank you. [01:25:41] **Tsinghua University Representative**: Okay. Thank you. [01:25:44] **Peng Liu**: Thank you. So we move to next. [01:26:04] **Kehan Wang**: Thank you. [01:26:05] **Luis M. Contreras**: So hello, everyone. This is Luis from Telefonica. I will provide a refresher on on on this draft that we were presenting in in other ITF meetings, which is about the usage of a computer over traffic steering for mithole networks. So the scenario we are referring to is the the one that you can see depicted in in the figures. So, basically, ORAN, what is defining is the functional radio split, the functional split of the of the of the radio. So the the previous capability that were embedded in the business station is also how separated and divided between real time processing, non real time processing, and then the, yeah, basically, the the final process until the the the information is delivered to the network. So with this functional split, the the the capabilities of this DU, which is the the real part the real time part and the CUUP, which is the non real time part, are instantiated and distributed, and it could be distributed across different data centers. That is what you can see on your left on your right. Sorry. So what we are trying to address is the connectivity of those these parts, And, basically, are the red cycles or the red squares that you can see in the figure. So, basically, all these components from ORAN, the DUCU will be connected to to routers. It could be equivalent to the cash for orders. In the oral terminology, they are denominated transport network elements. And and then that we could use different encapsulations shell underneath certain automation six, MPLS, or whatever. So the point here is is basically to try and to define what could be the interplay between the management a couple of management systems in ORAN, which are defined as ORAN SMO, service management orchestrator, and the CAD entity so that we can basically accomplish this connectivity according to the indication of the Oran SMO. So going a little bit into the relationship between these management and control entities in Oran and ATF, there are this SMO, the the management entity that is highlighted there in in yellow in the figure, which basically would decide what are the the the endpoints to be connected, the DU with the CU, with the corresponding CU. Then we have the O Cloud, which is a management system of the virtual I virtualization infrastructure where these DUs and CUs are instantiated. And, basically, it's it's a cloud manager. It's highlighted in blue in the in the figure. And then we will have the the transport network manager that will give the scope of a ATF in this case, which basically will be the one enforcing the connectivity and performing the traffic steering at the end between the EU and CU. Here, in in terms of how this could be materialized in in terms of technologies in ITX, we are network less controller behind for enforcing the the the path, but the decision will be taken by the CALSI elements, the CALSI path selector so that we essentially could define the better path to be a steer for the traffic connecting the DUs and the CEUs. Just jumping back a little bit to the previous figure, as you can see, the DU could be connected to several CEUs. So the point here is to decide the proper one according to the network metrics and the service metrics and cloud metrics that we could have at every moment. At the moment, I'm taking the decision, I I mean. So going moving forward in in in into this, so the the there is a gap in how the ORAN SMO and CALS can interface together. One presented before the the data model for CALS that that they work in that other graph. But, basically, that that work is helping us to to enforce, let's say, the decisions that are provided by SMO, but this somehow in in that there is missing parts, like the the how how to do do the interplay between the SMO and the card's entities. So now in Norand, the SMO architecture is is transitioning from the classical one where we could expect from interface on API to be consumed from the transport network to some other architecture where they are defining two entities. Basically, the data exchange the data management and exposure entity that is basically used for exchanging data between external management systems and the service management exposure entity, which is basically to leverage on the services that are offered by other components as could be the the transport ones, the ones solved by the ITF site. So all these two entities will facilitate the the interplay between SMO and and and T and M, so all the CAS entities. So our work or proposition here is to keep working on that and basically trying to define properly so that we could have the the proper interplay within ORAN world and ITF world. So next steps that we see for this effort to describe the architectural implications of the interworking with an ORAN SMO and and CATS. So in summary, analyze how the CATS entity can be connected to both the DME and the SME in the ORAN side, and then basically explore other scenarios like the dynamic changes that could come due to service affinity, interaction, or even issues in the network. We will start keep working on that. We will prepare new version for next ITF. And and as I basically, the idea here would be to be consistent with whatever is being produced by ORAN in the working group line, which is the working group devoted to transport to to transport technologies. That's all from my side. Thank you. [01:31:45] **Peng Liu**: Yes. Then Carlos. [01:31:48] **Carlos J. Bernardos**: Yeah. Carlos Bernardo, Just one comment. I think this may be related to the DMM draft on traffic steering, mobile traffic steering. You know, with the one from Marco and [01:32:00] **Luis M. Contreras**: Could yeah. I I know what what what draft that you are referring to. [01:32:04] **Carlos J. Bernardos**: Just a comment too. Maybe you you can Yeah. Take a look because I think there are some potential relationships in [01:32:09] **Luis M. Contreras**: Thanks for for the pointer. I think that the ambition of Marco's draft goes a little bit beyond our staff of connecting with the, what is calling in three PP side the data network and so. But, yeah, there could be similarities. I will look at it. Thank you. [01:32:26] **Adrian Farrel**: Hi, Lewis. It's Adrian. I'm struggling a little bit with an architectural view here as to whether this is a layered architecture or a stitched architecture. So maybe go back one more picture, that one. Are we traveling across the ORAN and then stitching into an IP network, or are we layering ORAN on top of an IP network? [01:32:58] **Luis M. Contreras**: Yeah. I will say that this how the the the ramp up, we kind of overlay in the sense that the I mean I mean, we'll go we'll be more in the in the direction of the stitching in my view so that basically, we provide the connectivity to some entities that are governed by the the SMO. So the SMO assumes that the connect that the transport will be there, but similarly to the point in as three VP does. And and, basically, we will need to stitch the the the different entities for serving the connectivity purposes. [01:33:39] **Juan Deng**: From. I noticed that you mentioned that that the existing data model is not enough for your scenarios. So do you have some suggestions for the more information to be carried in the data model, or you have another document planning? [01:33:58] **Luis M. Contreras**: No. We want precisely to work on on on that, to understanding where the gaps so that probably some parts could could be even moved to the data model draft that you presented before. But, basically, in our understanding as today is your the the draft of the data model draft is somehow internal to the to the CATCH part, but not not not working on how to interface the CAS part with something external. So that would be the part that we could cover here. [01:34:26] **Juan Deng**: Yeah. Okay. If you if you have the general information to be carried in that model, where can for your con contribution? [01:34:38] **Peng Liu**: Yep. Yeah. Thanks. [01:34:39] **Luis M. Contreras**: Of course. Thank you. [01:34:44] **Peng Liu**: Okay. Thank you. [01:34:58] **OTN Presenter**: Okay. Good afternoon, everyone. I'm pleased to present our draft framework and applicability of cuts in optical transport networks, which is now in o one revision. In this session, we will explore how the CAS framework can be extended and supported in the OTN scenario. First, let's establish the scope of this draft. As we know, the CAS framework is versatile and technology neutral. Our draft doesn't propose a new framework or modify the core CAS model. Instead, we complement it. We extended the CAS awareness into the optical transport domain as a scenario extension. Our main objective is to provide an additional underlay network option. So high performance sensitive workloads require something more deterministic. So by integrating the o 10 characteristics with the compute with the computing layer matrix, we can make a joint multidimensional steering decisions. So to keep our work clean and aligned, our draft focuses on the on how the CAS functional components, such as CSMA, CNMA, CPS, and CTC operate and cooperate over an OTN infrastructure. So why specifically target OTN and FG OTN forecast? It comes to two requirements, the hard isolation and deterministic performance. The high performance workloads are sensitive to the packet loss and jitter. By matching the service flow into the optical containers, we can provide physical isolation. It can eliminate the noisy neighbor in the transport layer. And furthermore, the OTN can provide deterministic performance. It can guarantee the sustained high bandwidth and deterministic latency. When you are synchronizing parameters across massive distributed clusters, the unpredictable latency may cause the GPU to idle. And the the OTN underlay ensures that the transport network is near the bottleneck. So let's dial into the new updates, which is detailed in section four dot five, distributed accelerator assisted computer services. In real world deployments, the operators usually are placing the accelerator clusters across the geographically distributed sites. And the available capacity at each CS each site fluctuates dynamically based on the local workloads. To address these distributed accelerator requests, we are facing two requirements. So compute aware and network aware. Steering based on only one metrics is insufficient. If we only look at the compute metrics, we may route the traffic through a congested network. And if we only look at the network metrics, we may end up at a site with zero GPU availability. So by using the OTN forecast, the CPS can jointly evaluate both domains, selecting the optimal site while simultaneously establishing a deterministic optical pipe. Let's walk through the CAS aware OTN workflow. First, the matrix distribution. The CSMA monitors real time accelerator availability, and there's a CNMA gathers available audio k or FGOT in bandwidth and optical latency. And we can also use the PCEP LOS as the protocol. Second, the classification and the joint selection. The ingress cast traffic classifier identifies a computer request, and the the CPS selects the optimal target instance and computes the optical path. Third, deterministic transport. The ingress edge node encapsulates the flow into a dedicated OGUK or FGUK container. So, in the result, it is highly reliable, hard isolated pipeline. It can provide a proactive performance assurance from the first byte to the task completion. To wrap up, we believe this draft provides an additional underlay option that demonstrates how the CAS framework can be applied to the optical transport networks. Moving forward, our next steps are, first, to align the applicability analysis with the official OTN official CAS use cases. And second, conduct comprehensive gap analysis on the existing protocols and the data models to identify what needed to extend. And we formally request for working group adoption for this work, and we welcome your comment, feedback, and the collaborations. Thank you. [01:41:29] **Peng Liu**: Oh, thank you. Nobody in the queue. Okay. Thank you. [01:41:36] **Kehan Wang**: So [01:41:38] **Peng Liu**: we move to next. [01:41:45] **Adrian Farrel**: Yes. Longfei, it's your lucky day. [01:41:57] **Longfei Xia**: Yes. Thank you. Thank you so much for arranging this schedule for me. I [01:42:02] **Peng Liu**: give the control to you. [01:42:05] **Longfei Xia**: Okay. Okay. I can I can see the Slack clicker now? And I hope everyone can hear me. Is the the voice okay? Yeah. [01:42:12] **Adrian Farrel**: Can you try to be loud? [01:42:14] **Longfei Xia**: Okay. Is the voice okay for now? [01:42:17] **Adrian Farrel**: Better. Better. Thank you. [01:42:19] **Longfei Xia**: Okay. Thank you. So hello, everyone. This is from Channel Mobile, and sorry for joining this sorry for joining online today. So what I would like to share with you is my proposal on operational semantics for cast metric cast metrics consumption. The issue we are looking at is a quite simple issue. A metric may arrive correctly at the decision point, but the condition behind that metrics may already have changed. The update per size may also have been interrupt or different sources may now provide different views. So the metrics is still available, but whether it is still suitable for steering, selection, admission, or other resource decision is no longer obvious. All we would like to discuss with the working group is a common way to describe three things. Whether the metric is still suitable for the intended decision, how the consumer should handle it, and whether that handling is visible to the operators. Okay. SSI changed because I have already clicked it. [01:43:22] **Peng Liu**: Yes. It's change. [01:43:25] **Longfei Xia**: Okay. Thank you. So why do we need additional semantics there? Catmetric may describe computing load, serviceability, headroom, or other conditions that can change quite quickly. One option is to simply distribute every metric more frequently, but that increase signaling and the processing overhead. Frequent change may also make steering less stable. If we reduce the update frequency, the overhead come lower, but the consumer may then be make some decisions based on outdated view. So today, we would like to use useful metadata set we we already have some useful metadata such as timestamps, age, revisions, sequence number, and the validated information, but there are still pieces of evidence. They do not directly answer the questions the consumer side really cares about. Can I still use these metrics for some particular decision? The distributor system has already this has already deal with some similar problems through ideas such as bonded stillness and the uncertain intervals. For CAS, we would like to introduce a similar operational distinct distinction into metrics consumption so that we can distinguish normal use, constant use, and unsuitable use. Slides. Hopefully, it changed. And another important question is where the judgment should happen? Or we use that it should happen where the metric is actually consumed for our decision. Yeah. Centralized CPS, the main concern maybe the delay introducing the collection, transport, pro processing, and aggregation. By the time the metric is used, the globally collected view may already be old. These are ingress embedded decision function. Different ingress point may hold different versions, different source subset, or different update histories. In a hybrid in a hybrid deployment, central and local input may also differ in age, trust, scope, and the intended use. So the same metrics may require different consumption decision in in different deployment modes. Based on this gap, we we propose three operational semantics. The first one is freshness. It describe whether the metrics still reflect the underlying con condition with sufficient temporal confidence. This judgment does not have to rely only on age. It may also consider revision gaps after continuity validating information, uncertainty bound, and consistency across so across sources. The second one is operational acceptability. This describe what the consumer has allowed to do with the metrics for the current decision. The metrics may support normal use, and it may only support the and it may only support the reduced trust of cost gain or fallback use, and it may also need to be excluded entirely. The answer may also depend on the purpose. For example, a steel utilization value may still be acceptable for a cost fallback decision, but no longer for price admission control. The third one is assurance exposure. This makes the resulting states the reason and the actual handling visible. This that matter because forwarding may still appears normal from from the outside, but the has already entered a degradable mode. So together, these three semantics connect the available evidence, the resulting consumption decision, and operational visibility. So what does this draft added to the current CAT work? Existing CAT metrics definition describe what a metrics represent and how it may be carried on abstract. What is still missing is the consumption side, meaning when the metrics reach the decision function, is this still suitable for this particular use? I also have discussed this point with Cohen. It is still whether it is still suit one of the author of the catchment definition draft. And he suggests us the freshness could be incorporated into the draft as a operational consideration. And in this presentation, I would like to take the discussion one step further on the one step further and explore what does semantics may mean for the practical the protocol protocol operation OEM without without prescribing any specific protocol behaviors. From the protocol perspective, it means that we do not need to treat every metrics in the same way or simply increase the update frequency of everything. Different metrics can have different dynamics and different rates. For example, some may need frequent update and some may tolerate less frequent update. And some may still be usable in a debride mode of after crossing a soft bra soft bond, and others may need to be excluded or triggered by fallback after crossing a hard bond. From the OEM perspective, the operators should be able to see more than the final steering results. You should also know whether the result has been produced from normal input, degraded input, or fallback be or fallback behaviors. Otherwise, the handling of this stale or inconsistent inconsistent information remains hidden inside the implement implementation specific logic. And also from the metric schematics perspective, this draft does not replace the time stamp TTLs or something else or existing OEM mechanism. Instead, it gives source mechanism a common operational meaning for the metrics consumption. And this can allows the metrics with producers, distributors, consumers, and the operators OEMs to share the same understanding of when a metrics can be used normally, when it should be used with constraint and when it should no longer be used. And if time allows, I also have three very simple examples. Maybe this page. [01:49:32] **Adrian Farrel**: Yeah. Maybe pick one of your three examples. [01:49:35] **Longfei Xia**: Okay. Thank you so much. Just three examples are how it could work in practice, how the states could be divided, and what's the resulting, for example, Yam OEM or OEM information could be look like. The first example shows one possible real is hopefully, that is current page six. And this example shows one possible realization of three three semantics. First, the consumer develop a freshness state from the available evidence. That evidence may include, for example, observation time, update time, revision, sequent continuity, validated table, uncertainty bonds, or source con consistency. And next next, the consumer combines that state with source trust, the intent decision use, and the local risk bond. This produce the operational acceptability state. And a steel matrix is and a steel matrix is therefore not automatically unusable. It remains within a soft bond. It may still support our reduced trust or fallback use. And if it exceeds a hard bond, it may need to be exclude. Finally, the implementation may expose the resulting state, the degradation reason, and the actual handling through, for example, young operational state telemetry, OEM diagnostic, or troubleshooting record. And these examples shows a few possible ways to divide the freshness and operational acceptability. For example, the time or age based, uncertainty window based, update continuity based, or source trust based or consistency based. And for the final example, let me shows show a illustrative example to show whether to show how the proposed semantics could be applied. On the left, we use a YAML style format simply to illustrate the processing flow, and you can see we have three metrics. Each metric has its own evidence, freshness state, operational, acceptability state, and actually how and actual handling. For example, we take the first one as an example. The compute utilization metrics is older than its freshness found. Its freshness state is therefore becomes stale. However, it's still considered as a degrad rather than completely unusable. So the consumer can use it with reduced trust, and there's something similarly for the second and third metrics, for example, admission, headwear, and network latency. And on the right side, we have we show two possible way to expose the result. A young operational state could equal could e could equal equals the result and could expose the current state and the actual handle of each metrics. And OEM events record can report the state change, the reason, the resulting action, and overall that related condition. Okay. So that's all I want to share today, the draft's proposal, call mathematics, and the information that should be visible. We hope this proposal can start a discussion on the common operational semantics for CAS. And we will really appreciate your feedback on this problem statement and the proposed domestic and any possible direction of this draft. Thank you so much. [01:52:58] **Peng Liu**: Thank you. Thanks for your work. Anything you want to see? [01:53:04] **Adrian Farrel**: I was just wondering whether Kehan wanted to say anything. No, don't join the queue. Just yeah. [01:53:12] **Kehan Wang**: Thank you, Yes. I've talked to the author of this draft. I think it provides some useful input on operational considerations into CAS metrics. But I think beyond that, it also touches, like, security issues. So yeah. So if you're willing to, you can just raise PRs or send mails to the authors of Casmetrics draft and before the working group last call. Thank you. [01:53:38] **Adrian Farrel**: Thanks. Yeah. And and thank you. I'm this draft arrived a little late, but I really like the work. So I I hope we can achieve something with it. Alright. That brings us to the end of our agenda. I think we've got two things to cover from the chairs. The first is to repeat this message about protocol work. Do not go to other working groups claiming that you need a piece of protocol because cats says it's important. Come to cats, explain to us what the protocol requirements are. We can agree the protocol requirements and then we can work with the relevant protocol groups. That's the first thing. The second thing is at the start of the meeting, did a rather feeble attempt to get a feeling for who amongst all the cats will be going to San Francisco and who will not. Neither of the chairs will be there. We would be willing to chair a meeting remotely, but we have a feeling that a lot of people won't travel. In which case, we should do maybe two interims, one each side of San Francisco to cover the material perhaps in more convenient time zones. So if you have opinions on that, drop an email either to the list or direct to the chairs. [01:55:16] **Kehan Wang**: Okay. [01:55:19] **Peng Liu**: Thank you. We have five minutes. Anyone has any questions? Want to talk? Okay. So let's close the meeting. Thanks, everyone. Thank you. [01:56:36] **Luis M. Contreras**: Let's go. [01:56:42] **Kehan Wang**: Oh.