Markdown Version

Session Date/Time: 24 Jul 2026 12:00

[00:02:14] Moe: Shocker. Friday after lunch. I'm just gonna give it just another minute or two.

[00:02:26] Jonathan Lennox: There's people in the room, I see.

[00:02:28] Moe: Yeah. The presenters are all remote, and Greg is remote. So should have at least six or something online. Or oh, I see thirteen online. And oddly enough, the clock here still says it's two before three minutes before two. So, hopefully, that's not the same clocks outside too. Just give it one more minute, and then we'll get started anyway. So for

[00:05:17] Jonathan Lennox: some reason, that says resume is fourteen. This says fifteen.

[00:05:24] Moe: Mine says

[00:05:35] Jonathan Lennox: Oh, maybe one is counting slides clicker and the other one is

[00:05:40] Moe: The slide clicker is a user? I see 15. and MeetEcho is for Zulip.

[00:05:48] Jonathan Lennox: Yeah. And participants cancel slide clicker.

[00:05:51] Moe: K. Yes. So I just got one slide update that I that request that I'm trying to upload. Okay. Sorry about that. Just had to get one final slide update in. Okay. So we're gonna have a quiet session today because that's what's typical for Friday afternoon after lunch. Welcome to the ML codec session at IETF 120 in Vienna. So the IETF Note Well has three main pillars, the anti harassment policy, the intellectual property rights policy, and the privacy statement. But just to recap those, you you're expected to disclose any IPR that you're aware of, either yours or others, by any participation in any IETF activity. You agree to abide by all these rules. And just some meeting tips. I think for the only few in person meeting participants, please make sure to log in to the Meetecho on-site tool so that we can get your attendance for the blue sheets and so that you can raise your hand and join the queue. And the resources, if you have any problems with Meetecho, are right there. The full agenda is available there. Our agenda is here, and the notes are at that URL. I don't think we need a notetaker. I think we'll rely on the automated notes. So I won't ask for anyone to waste their time spending transcribing anything. Agenda today, we have the Opus extension mechanism from Tim, deep redundancy from Jean Marc, speech coding enhancements from Jan, scalable quality extension from Jean Marc, and then the testing battery from Camille. Any agenda bash or anything people wanna mention about this agenda? Alright. And hopefully, we will end early today. So first, let me just cap off what the current state of the working group documents are. We have an OPUS extension. We we finished a working group last call, but we need to start a a short, quick new one. And we hope to update the milestones to get this to the ISG before the next meeting, so by by November. Dread, Jean Marc has a request that he thinks the material's mostly done, so let's just confirm with the work group. And if so, then maybe we'll progress that to our group last call and also try to get that to the ISG before the next meeting. And then for speech coding enhancements, we'll let Jan comment about what he thinks the maturity of it is and when when he thinks that would be ready to progress. But we're hoping by by two meetings from now, so March 27. And the same for the OPUS HD scalable quality extension. And the test battery, we don't have a don't have a milestone for that yet, so we'll we'll need to add one. But we'll work with Laura and Camille to figure out what the what the best date for that would be. Okay. So first up is Tim to give a quick update on the, Opus extension.

[00:11:26] Jonathan Lennox: Alright. Thank you, Moe. This should be very quick. Next slide, please. So the update is nothing changed. There's a new version just to avoid the draft expiry, but next slide. And I think we're ready for the second working group last call anytime the chairs are ready. And that's it.

[00:11:53] Moe: Yep. Thank you, Tim. We'll we'll the chairs dropped the ball on this. Last time we had the first working group last call, there were a lot of comments from from Roman, and I think Tim addressed all of them. I think Roman confirmed that all of his issues were addressed, and there was a new draft revision updated in in December, I believe. But the chairs forgot to. We overlooked progressing it immediately so that so that it would get frozen. But because we didn't do that, it expired. And and now this refresh here is just to start a quick second work group last call to make sure that there's no more comments and that the changes that were made for the previous comments are agreeable to everyone, not just Roman. But then, hopefully, we'll just do a a short two week second last call and then start the Shepherd ride up after that, barring no substantive changes. Alright. Thanks, Tim. And thanks for your patience on this. That's No worries. Taking longer than expected. Alright. Moving right along. Jean Marc, you want me to share or you or you wanna share?

[00:13:04] Jean Marc: Please share.

[00:13:05] Moe: Okay. Would you like remote slide control, or do you want me to flip?

[00:13:15] Jean Marc: I don't know how the remote slide control works, so probably, yeah, I should flip.

[00:13:22] Moe: I'll just flip for you. Just tell me to flip.

[00:13:24] Jean Marc: Yep. Okay. So, well, this is the update for deep redundancy.

[00:13:30] Jan: Next slide.

[00:13:34] Jean Marc: So first, what changed since since the last meeting we had in IETF 119? Essentially, on the model side, there was no change no change to the model, no change that breaks compatibility or anything. There's been a few fixes to the implementation, most notably how the offsets are well, the offset was correct, but the decoder implementation would not take it into account correctly. It didn't break compatibility, but it could leave you with fewer DRED frames that you had that was fixed in the implementation. At this point, I believe that the DRED description in the draft is complete, but I would very much like feedback there. And the main thing is I've update I've uploaded what I believe is all of the DRED material that we would want at the URL I have on the slide. This would be obviously a temporary location. I believe the ITF has found the right place to put all of this material.

[00:14:51] Moe: Yeah. So, unfortunately, our AD, our new AD, Charles Eckel, isn't here today. He he had to leave the meeting early. But I believe we have identified the permanent arch archival home for this, and the only question was about the total size. These materials are the are all the only small, like, a few few megabytes of materials, not the gigabytes of of training data. Right?

[00:15:18] Jean Marc: No. I I've excluded the training data from there. Okay. It's kind of insane.

[00:15:23] Moe: Okay.

[00:15:24] Jean Marc: Yeah. Plus it's downloadable from existing sources anyway. I'll I'll go through the material in the other slides anyway.

[00:15:33] Moe: Jonathan, do you have a question?

[00:15:34] Jan: What is the current location, sir?

[00:15:37] Moe: The actual URL? Yeah. I'm sorry. I I forget what it that's an official IANA archival location for Blobs. And they were okay with it as long as it was in, you know, manageable, you know, small gigabytes or, you know, or fractional gigabyte range. But, you know, many tens or hundreds of gigabytes would would not be accepted. Okay.

[00:16:12] Jean Marc: So this is just a reminder, for test vectors. I've presented this slide in the past, but, basically, there's the test vectors are all all in a big tarball. There's some scripts to check them. And in terms of coverage, they will check they will check the feature decoding. There is a very strict criteria on there. There is a check for the Vocoder that is pretty loose that just checks that whatever gets synthesized from features by the Vocoder kinda makes sense. And then there's a test for the full DRED integration in Opus, and that one is not normative. It's just if you have a very similar implementation, you could tell you if you've got a bug or not. But that's how the test vectors are done.

[00:17:07] Moe: And just to clarify, is the Vocoder only you say normative here, but normative for the testing? Just normative for the test vectors?

[00:17:15] Jean Marc: It's normative in the sense that there is something in the draft that says you must do that. Like, the Vocoder must satisfy these conditions, but the test is meant to be pretty loose. Essentially, the the Vocoder test essentially defines the features themselves. Like, we have these acoustic features, and you we haven't defined how they are computed. We've given a a recipe, but the computation is not normative. Therefore, you need the Vocoder to sort of obey the the pitch and the the spectral envelope that's encoded in these features. Like, if you if you make a Vocoder that has, like, twice the pitch, then it means you've disregarded the you've not interpreted the pitch feature correctly. That's basically how it is.

[00:18:17] Camille: Got it.

[00:18:22] Jean Marc: So next next slide, please. Okay. So this goes through the the artifacts I put in the temporary directory. There's, I believe, eight of them. Three are normative. So first one is the model weights. They're published in this file, dread unquantized weights dot bin. The format of this binary file is defined in the in the draft. Then there is the test vectors that I mentioned in the previous slide. And there's a comparison tool, and there's three c files in a shell script to do the actual comparison. I think it makes sense to have those in these artifacts. I could have a copy in the draft. I don't have a strong opinion there. It's, like, 100 lines of code. It's not huge, but it's not just 10 lines there. So that's for the normative artifacts.

[00:19:34] Moe: And and no, I guess that wouldn't

[00:19:45] Jean Marc: go ahead. That's

[00:19:46] Moe: No. No. No. I was I was I was trying to figure something out for for the way so the the these files, I assume you you wanna add some integrity check, some hash or something for them so that you have confidence that what you produced is what actually gets posted?

[00:20:07] Jean Marc: Yeah. It may actually make sense to have so you you you're talking about having something like a shot to 56

[00:20:13] Moe: Yeah.

[00:20:13] Jean Marc: Hash in the in the in the draft

[00:20:19] Moe: that we can verify? I don't know about in the draft, but at least when you hand off something, you know, I would like you to generate the hashes of what the files that you think you handed off are, and then we'll preserve those whether we put them in the draft or or just as another file in the archive. I guess it wouldn't hurt to put them in the draft if we knew.

[00:20:43] Jean Marc: Yeah. I was going

[00:20:44] Timothy B. Terriberry: to suggest putting them in the draft. I think that might make sense.

[00:20:47] Jean Marc: Yeah. Because they're immutable in the in the draft, so whatever happens, it can be checked.

[00:20:52] Moe: Yeah. That's

[00:20:52] Jean Marc: I think that makes sense.

[00:20:54] Moe: Yep. That's a good that's a good good check on on the archive being sane. Okay. So let's let's do that in in an update if you're confident these aren't gonna change.

[00:21:11] Jean Marc: I mean, if they change, I'll change the draft. Right?

[00:21:14] Moe: Yeah.

[00:21:15] Jean Marc: But, I mean, like, once it becomes RFC, no. They're not they're not going to change. And if they would change, we would need to publish an updated RFC.

[00:21:26] Moe: And when you mentioned that no model changes, do you mean exactly zero changes to all the weights too? Or you just mean

[00:21:33] Jean Marc: No changes the weights model. No. No. No changes to the weights. Okay. There there were, like, a few changes to the c code that were, like, not normative.

[00:21:44] Moe: Yeah.

[00:21:51] Jan: Okay.

[00:21:52] Jean Marc: And so now the informative artifacts, there's five of them. First, the Python model. It's got, again, all of the weights in there for DRED. It means if you wanna play around with the model in Python, you can do it. You can verify how it's been converted to the blobs. That being said, at some point, PyTorch is probably gonna break compatibility, making it harder. And at some point, they're gonna be mostly useless, but I think it's still useful to have them in the informative artifacts there. Then there's a blob with all of the weights. So complete unquantized weights, that means all of the weights from the the Vocoder, which is not normative. All of the weights also from the packet loss concealment, you can have, like, a full implementation. So that's why they're in the informative artifacts. And then there's the quantized version of all of these weights in there. The complete c reference implementation is also there, and all of the Python scripts that were used for the DRED model or the informative artifacts. So at this point, aside from the training data, it is basically everything I've had or used to to come to to work on DRED.

[00:23:40] Jonathan Lennox: Jonathan Lennox, this isn't really a working group thing, but it just is a question for you. Do you anticipate that the the reference implementation would stay exactly this, or do you wanna do an actual Opus release matching the RFC so there's a actual version number that's not a GitHub snapshot?

[00:23:58] Jean Marc: That's a good question. I mean, we so there there's kind of a chicken egg thing there in that once this is published, we would probably once this becomes becomes RFC, we would probably do a 1.2 release, but I wouldn't wanna do a 1.2 release before this gets accepted. So I don't think that can be an I don't think that can be a real Opus release, especially considering the fact that the one thing that we'll need to change in the final c reference implementation, which is not done now, is to change the extension ID to be the final extension ID. And that I will not put in a release before this thing is officially

[00:24:46] Jonathan Lennox: I mean, doing this, like, at RFC editor time. You know, once you have a RFC number and have a note saying, from note to RFC editor, this one hash will change when we do the final OPUS release, only just give the RFC number and the whatnot. And we'll get the so, I mean, I don't think do it now, but I think, you know, like, as we get through the RFC process, having basically effectively having the Opus release and the RFC be effectively a cluster with mutually referencing each other. You'd have to coordinate with the RFC editor and probably talk to them, but I don't think it'd be too hard.

[00:25:22] Jean Marc: Yeah.

[00:25:25] Jonathan Lennox: I mean, it's just a thought. Doesn't have to be.

[00:25:27] Jean Marc: No. Yeah. I'll I'll think about it, see if there's, like, a sane way of doing it. Thanks.

[00:25:35] Moe: And just to clarify, none none of the artifacts are any in any way dependent on what's in the OPUS repo and what gets snapshot at any time in the OPUS repo. Right? These artifacts are everything that we submit as artifacts are gonna come directly from you, integrity checked, and then passed on to IANA. Right? We're not pointing anyone to an OPUS repo any repo anywhere. Right?

[00:26:03] Jean Marc: No. It it's all like, the c reference implementation, it would be able to use all of the binary files.

[00:26:09] Moe: Okay. So John so Jonathan's point is just about integrating the integration of this into Opus, having a release snapshot that points to the the real RFC version of this just for convenience of developers to know that this is the RFC version of Opus with DRED. Right?

[00:26:31] Jean Marc: Yeah. I mean, I I guess it would be it would be about making if my understanding was correct, it would be making the c, the complete c reference implementation be an actual Opus release is what was suggested as opposed to a particular Git commit or something. That's what you meant, Jonathan. Right?

[00:26:54] Jonathan Lennox: Yeah. Yeah. I mean, act act both directions, really. Basically, having the c reference implementation be an actual gop OPUS release and that being an actual OPUS release, which is that that also able to cite the RFC number. So it just it means that these have to be released Coordinate. Basically as a cluster, you know, as it were from the RFC editor's point of view. Basically

[00:27:17] Jan: Yeah.

[00:27:18] Jonathan Lennox: The RFC editor assigns a number. We generate the final OPUS release with the RFC number in the code, publish both at the same time.

[00:27:28] Timothy B. Terriberry: I I would also throw out ignoring even the OPUS release itself. It would be a little strange that the reference implementation were still using the experimental ID.

[00:27:37] Jonathan Lennox: So it would be Definitely. But but that that probably happen. I suspect that that would be a the I mean, the IANA ideas assignment happens before RFC number assignment. That would, I think, probably happen at around, like, working group last call time or a little after.

[00:27:55] Timothy B. Terriberry: Yeah. Because that's one thing we point is wanna It's just that's going to create this kind of coordination a little bit of coordination juggling too if the reference implementation Yeah. Is using the correct ID,

[00:28:05] Jan: which I

[00:28:05] Jonathan Lennox: think So we need to figure out that's actually a good question. When do we wanna do those ID assignments?

[00:28:09] Jean Marc: Yeah. I mean, essentially because to be able to use that ID, you need two things. One, obviously, is to have the IN assign it. Mhmm. But most importantly is to know that you're not going to break compatibility.

[00:28:27] Jan: Right. Right.

[00:28:28] Jean Marc: If you release something using that number and then someone's, yeah, we need to change something and you break compatibility, then

[00:28:33] Jonathan Lennox: that's Yeah. So, I mean, I think it sounds like that's what I'm saying. The working group needs to decide when we're when we're saying, yes. We're going to the real number now.

[00:28:43] Moe: I think when when IANA does the code point allocation and then, you know, announces it to the authors Mhmm. Then at that point, I think it would be safe to update the code to use that number. And then if the code also needs to have a comment about the RFC number, then we'd have to wait for the RFC editor to assign an RFC number. I'm not convinced that the that the RFC number needs to be in the anywhere in the Git repo. Could you say it's standard, you know, standard DRED?

[00:29:17] Jan: Yeah.

[00:29:19] Jonathan Lennox: Yeah. That's that's true.

[00:29:20] Moe: Yeah. Really need is the code point for for the

[00:29:22] Jonathan Lennox: So I mean, we need

[00:29:23] Moe: to So for the extension.

[00:29:24] Jonathan Lennox: Yeah. Exactly. So we need to figure out when we ask IANA to do that registration, which I think which I would suggest probably at time of and I'm not sure when does what what is when does IANA normally do final you know, we we could ask for early registration if we want to Yeah. As soon as the I guess, soon as the registry exists, which is publishing the first draft. But do we want to, or do we wanna do it? I guess, normally, it's after ISG review.

[00:29:58] Moe: I I would like to wait until after ISG review because we don't know what we're gonna get from the ISG. So I would not wanna do if the process was gonna be get the code point and then update the code, I would not wanna update the code with the code point until we had confidence that that was gonna be the code point and the implementation was stable and wasn't gonna rev.

[00:30:17] Jean Marc: And just just to clarify, so once it is approved so at some point, IESG approves, and then there's some lag before, like, RC editor and auth auth 48. But, essentially, from the point where the IESG approves, we can pretty much assume that at that point, we're not breaking compatibility. Right?

[00:30:41] Moe: There should be no substantive changes after IESG review.

[00:30:45] Jean Marc: Okay. So I guess at that point changes even be safe to make an Opus release.

[00:30:50] Moe: Yeah. And, actually, it's it's probably not even contingent on this draft. It's probably contingent on the Opus extension mechanism draft. When we have confidence that that got through the IISG without any changes because that's really what would impact the code point allocation. If IISG's review of that said, oh, no. This is stupid. You need to do this. And then our code points are, you know, a totally different, you know, space or or something different about the design. That's what would impact this this code point allocation more than this actual draft. So I think once we get extension mechanism progressed and through IESG, they will have a lot more confidence that that just assigning the code point number would be will be easy in a slam dunk. Yeah.

[00:31:29] Jean Marc: But it's more about breaking DRED compatibility. So it seems to me like we could even cut like, from the right at the point where IESG approves, we would be able to make an Opus release anytime from that point and then just reference that release in the in the artifacts. Right?

[00:31:49] Moe: Yes.

[00:31:51] Jean Marc: Okay. Thanks. Yeah. I think there's Yan in the queue.

[00:31:54] Jonathan Lennox: And and and that is the point at which IANA normally does code plate assignment. So when once I issue approval is done.

[00:32:02] Jan: Great.

[00:32:06] Jean Marc: Go ahead, Yan. Oh, we can't hear you. Still can't hear you.

[00:32:16] Moe: We can see you, but we can't hear you. You do not appear muted on MeetEcho, so it must be some local local device mute. So I will wait for Jan to come back. Is there any objection to going to work group last call? The editors believe that it's stable and ready for progressing. I think we have Jan back now. You wanna try speaking again, Jan? Still cannot hear you. You appear fine on MeetEcho.

[00:33:31] Jonathan Lennox: If you wanna type something into the chat, I can say it at the mic if that's helpful. Yeah.

[00:33:51] Moe: So whatever you're typing may work for your comment about Dread, will not work for the next session, next presentation. Or I will do my best to channel you if you want on your slides. Okay. Yeah. It says he's switching to a different computer. Let's see if that helps. Okay. So for the notes, I'm hearing no objections to progressing this to our group last call. So we'll we'll start that shortly. Confirm on the list.

[00:34:37] Jean Marc: Maybe we can swap the order of the other presentations to

[00:34:42] Moe: get Sure. Let's do that. Jan, if you can hear us, we're we're going we're going to skip yours for now and come back to you. Alright. So Jean Marc, you're up again. Quality extension?

[00:34:57] Jean Marc: Yep. So now this is the update for the scalable quality extension for OPUS, also known as OPUS HD. Next slide, please. So just, again, a quick reminder of what this is about. This is about lifting some of the limits of the OPUS bitstream. So supporting more than eight bits of depth per band, supporting beyond 20 kilohertz of audio bandwidth and bit rates beyond or or up to five about 500 kilobits per second per channel. And do all of this while keeping forward and backward compatibility with RFC sixty seven sixteen. Next slide, please. So changes since the last meeting we've had, there's essentially been three changes that actually affect the spec. One of those is that we now allow allocation up to 14 extra bits for all of the bands. Before, it was just above 20 k. Now this is over the entire thing. It doesn't have a significant change, but it does break compatibility to signal that. There's also been, unfortunately, two typos in the bid allocation that were find found by LLMs. Pretty silly typos, but they actually it it was code that was doing the something that was very wrong, but in the end didn't change quality at all. But, again, it breaks compatibility to change that, so those two things were fixed in the code. I believe at this point that we have a complete description of OPUS HD in the draft. So it's it's not it's defined at the same level as RFC sixty seven sixteen. So it describes everything that is done, but we're still considering the code to be the normative thing here. The test vectors have been updated due to the bugs that were found. And last thing, recently, the Japanese audio society has certified OPUS HD as what they call a high res audio wireless. So, that's basically the update into a change.

[00:37:48] Moe: Curious what is curious what the wireless aspect has anything to do with this?

[00:37:55] Jean Marc: Oh, wireless as in these codecs are typically used for wireless because if you're gonna do high res wired, you might as well just transmit the thing over the wire. Right?

[00:38:08] Timothy B. Terriberry: Is is this a product of some conformance test or something like that?

[00:38:13] Jean Marc: Yeah. There's already a bunch of codecs that have this certification, I believe, like, l l dac, l can't remember the name. L l c three plus has this Yan still has PTSD over that. So but yeah. That that's the yeah.

[00:38:43] Moe: So so you you just submitted this individually as for certification to this body?

[00:38:49] Jean Marc: This was submitted as sort of as Google proposal to this it's not really a standards body. It's more like they're

[00:38:57] Moe: just Certification body. Yeah. Exactly. So this was branded more as a Google contribution a Google technology, not as an ITF product. Right?

[00:39:11] Jean Marc: In terms of the certification, yes. But the codec itself would have the certification.

[00:39:16] Moe: Okay.

[00:39:19] Jean Marc: Yeah. Okay. So in terms of conformance, and again, I presented this last time, the test vectors were just updated to match the changes. But they're the test vectors are all 96 kilohertz floating point audio. We've got three synthetic clips with a sine sweep, noise, and pulses. There's a speech clip, two music clips where the music was synthetically extended to 96 kilohertz. There's so there's a this time, there's a very strict decoder conformance test. The encoder itself is not normative, but we include an optional test if you wanna verify that, you know, the encoder actually does make sense.

[00:40:18] Moe: Just curious of was there any chance to get any true audio sources, capture ultrasound, something just to see if a captured versus a synthetic signal would make a difference? Try to capture ultrasound?

[00:40:34] Jean Marc: I'm open to anyone providing this, but I could not find real 96 recordings that were of very high quality with, like, good music and all that. Otherwise, it had, like, all kinds of noise and stuff, which wasn't great.

[00:40:56] Timothy B. Terriberry: Is I guess, like, redistribution terms also might be a a concern there. Yep. So I can't imagine you couldn't find anything at all.

[00:41:06] Jean Marc: Yeah. I found some stuff on, I can't remember, like, on archive.org or something that was c c zero or something. But, yeah, the the set of things you can choose from is very narrow, especially if you look at something that has good quality.

[00:41:29] Moe: Okay. If possible, it'd be great if you can expand the search to or if anybody in the work group can try to find a source that's true captured instead of synthetic, that would be really useful, I think.

[00:41:47] Jean Marc: Okay. So next next slide, please. So in terms of status, so everything is still on the main branch. It is off by default. You need enable QEXT and a runtime set QEXT to one to enable it. I've pointed here to the test vectors. The there's the QEXT compare tool that is working that works with the test vectors, and there's a fixed point implementation that verified as compliant passes all the test vectors, basically. So that's the current status. So the question is, is there anything anything needed, more to to be able to go to working group last call?

[00:42:54] Moe: Next slide. Was that a question to the working group or are you willing to advance?

[00:43:00] Jean Marc: That was question to the the working group, like, or any essentially, any thoughts here?

[00:43:12] Jonathan Lennox: Yeah. Don't thought it was you're saying that, again, the code is the is the normative spec. So unlike sixty seven sixteen here, we're planning to have this be in the or, you know, on online as well as rather than the ridiculous space 64 encoding in the document?

[00:43:33] Jean Marc: Yes. That was also part of I I assume that whatever we did for Dread, we can do for this. So have, like yes. Something not insane like base 64.

[00:43:46] Jonathan Lennox: Yeah. I mean, it would be nice if these two are actually the same release, but I don't know how easy that would be to do.

[00:43:53] Jean Marc: How what would be the same release?

[00:43:55] Jonathan Lennox: Basically, if the two if the the reference if if, you know, if the the snapshot the OPUS snapshot for both DRED and HD were the sexual same, like, OPUS release. But I don't know how easy that would be to do.

[00:44:11] Jean Marc: I mean, if if both are I think it's more of a schedule thing. Like, if both are considered ready at the same time, I'm pretty sure we can flag the the drafts that they're, you know, locked together or where where whatever.

[00:44:28] Jonathan Lennox: I guess, they'll just let the let the RC editor know know they're a cluster, basically.

[00:44:32] Jan: Yeah. Exactly.

[00:44:33] Jonathan Lennox: Don't explicitly reference each other. They would be a sort of a, you know, cluster with the release Opus release. So that might be a good idea. But, I mean, it just it that's not not necessary, obviously. It just it seems like if you have a snapshot of a note of the Opus tree as a informative reference in one and a snapshot of the Opus tree as a normative reference in the other, it'd be nice if they're actually the same code.

[00:44:56] Jean Marc: Yeah. I'm not against I mean, I don't think it's worth, like, delaying one by many months. But if they're around the same if they're gonna be published around the same time, that I would say it makes sense.

[00:45:14] Jan: Agree.

[00:45:17] Moe: Jan, you wanna try your audio again?

[00:45:20] Jan: I do. Can you hear me?

[00:45:22] Moe: Yes. I can hear you.

[00:45:23] Jan: Oh, wow. So I would have been out of out of ideas. Yep. I I just wanted to point out this actually the same thing Jonathan said. So there is also a dependency of the speech codec enhancement draft now on some reference implementation of a comparison metric, it would be great to also have that in one snapshot or that that could be referenced. So it's it's not on main yet, but basically would be great if that could go into the DRED reference snapshot so that it's out there.

[00:46:03] Jean Marc: In that case, I guess, like, it would be essentially, like, Office two release would be locked in with, like, all three drafts.

[00:46:12] Jan: What makes sense?

[00:46:15] Moe: Yep. Whatever makes sense on timing. I think, you know, if they all are pretty closely aligned in timing, I think it makes sense to have them as a cluster and and do a single Opus release. If they're, you know, drifting further apart, then it, you know, probably doesn't make sense to hold up DRED for anything. And it seems like that's the base anyway. The the comparison tool is in DRED. Right? So you need that to go to go out first. You know? And having a couple of snapshots is not gonna be a big deal.

[00:46:48] Jean Marc: I mean, technically so the comparison tool is different for DRED and Opus HD. So, technically, they could all go in different orders. I think DRED right now is the most mature one, but at the same time, QEXT is a lot simpler. So

[00:47:06] Moe: Mhmm. Yeah. So that that that's the main thing I wanted to ask the working group is I I think I heard your request, John Mark, that you you wanna progress the OPUS HD work now as well. You think it's ready for your last call as well. It's fairly new. I mean, it's only had two revisions. But like you said, it is much simpler. It's relatively straightforward extension of of a few OPUS parameters. Not a lot of new coding modes. So I don't think that from Chair's point of view, I I I don't see a process reason to to require more iteration. If the work group feels like it's stable and won't have any any new churn, I think we're we're okay to progress it as well. Does anyone see a reason not does anyone see a reason to hold off on progressing the Opus HD work right now? Okay. So hearing none, I think I think we'll be able to do a more group last call on on on this as well.

[00:48:23] Jean Marc: Great. So, yeah, this was I think, yeah, we can go now.

[00:48:51] Jan: Yeah. So hi. I will present about OPUS speech coding today. The focus is towards draft completion. Yeah. Next slide, please. A quick recap. It's been a while. So, basically, the story is we have a couple of enhancement algorithms in OPUS 1.6. Those break conformance. So we want to update the OPUS RFC to allow those kinds of enhancement methods provided it satisfies a set of requirements. I made an update from version three to four this week, actually, and that one is now kind of feature complete. Everything is in there. Next slide, please. So here's a quick overview. So the aim was to fill in all the missing pieces for a complete proposal. The qualification criteria, they are now complete for extending and non extending silk enhancements, and there is a new version of test vectors and explicit pass criteria. One of the bigger changes is I decided to drop enhancement with site information. We reported earlier about very early attempts that were not very promising, and we don't have anything to test. And I'm not pursuing it. So to me, it it doesn't make much sense to keep it in there just as as a hope for the future. That also means the new draft, it opts for STP for signaling, so to say, a little bit of information about the enhancement that a decoder plans to use. There are some security considerations now also added and a few missing missing details like prohibiting to train on the test data and things like that and mandating to do proper delay compensation before testing. There is also a paradigm shift. So earlier, I used I simulated batch enhancement methods by using barely trained versions of the models. What I did now is to actually distort the model weights by adding noise, which I think makes it easier for others to reproduce. There is also a Python script for that in the Opus repo. Next slide, please.

[00:51:30] Moe: Real quick, John, on this. So you're now abandoning any bitstream. There's no more side information?

[00:51:40] Jan: Yep. I mean, any side information would have gone into an extension. But since we don't know what that site information could be, I I'm not aware of anybody experimenting with this. I see little I mean, that that would be, so to say, just a blind guess.

[00:52:03] Moe: So this would be a purely decoder side

[00:52:06] Jan: Yes. Enhancement? It's purely decoder side and, yeah. The sole purpose of it is to allow this kind of decoder side enhancement.

[00:52:17] Moe: Okay. Yeah. Okay.

[00:52:21] Jan: So this is now an overview over the requirements that we have. So there are these extending think of it as planned bandwidth extension and the non extending algorithms, lace, no lace. So subjective evaluation, that didn't change. It's a should, that it should undergo subjective evaluation. There is this low band test that I presented before with the modified OPUS compare version that is a must for extending methods we or algorithms we now have also the high band test in place. That is a should. I will elaborate later why it's a should and not a must. Approximate face preservation, that didn't change. It's still a should. No training on test data is a must, and no training on the sources of source datasets of the test data is a should. So understanding that there is reason not to do it. But if somebody thinks they know what they're doing and carve out the part of the dataset that doesn't contain test data, then so be it. And interoperability, so to say, the statement that if you produce a bit stream that's decodable with an enhanced decoder, then it must also be decodable into a human recognizable reproduction of the signal with a legacy decoder. That's a must for both. Next slide, please.

[00:53:55] Moe: And just to clarify on this Yeah. Extending means extending to full band and high and high band test means full band test?

[00:54:04] Jan: It does extending does not mean you have to do full band. You could also do a super wide band, but the test is conducted at full band. So, basically, if you would only do super wide band, you would upsample to full band and do the comparison there. So the test is designed that should pass either way. It's it's not mandated that it has energy up to the highest band.

[00:54:34] Moe: Okay.

[00:54:40] Jan: Okay. So the low band test is basically the same as before. So the difference is the comparison metric is now implemented in c so that it can be referenced by the standard. There is the actually, interesting thing thing. So the version zero of the test vector failed even without enhancement. And the reason was interesting. So the 42 bit API changed the clipping behavior, and there was some clipping in the PLS SILK calc transition. And that basically made so few of the tests fail. That kind of highlights, so to say, the issues with fixing a reference decoder because a decoder can actually include a lot of non specified behavior. And the reaction was to do two things. So the first one is to normalize to a more conservative peak value. And then the second thing is that the comparison is now always against the reference decoder, and any reference decoder that is a conforming decoder is admissible. So if there are any changes regarding clipping or something like that or even resampling, then you can basically take a conforming decoder and the enhanced one and compare the two. That also solve elegantly solves the PLC problem that we've seen before or yes. Next slide, please. So the this is so to say the test results in blue are the is basically are the good methods, and in red is the distorted method. And as you can see, all I mean, 24, all the good methods are good. They pass all the tests at the at the given bit rate. And in red are the bad distorted models. And they typically fail, but at very low bit rates, it's kind of hard to make things worse. So there you have a few passes, but all would, so to say, fail the test that everything must test must pass. Okay. Next slide, please. There are no questions. The high band test is as presented at the IETF 120. The idea is that bandwidth extension should not be worse than doing nothing. So there is a per band distortion comparison between enhanced and, so to say, unenhanced and notes. Sorry. It's between enhanced and reference, and then between reference and low pass filtered signal. And, basically, you want the distortion between the extended one and the reference to be not worse than that between the reference and the low pass filtered one. That's also implemented in this osc_compare tool. That's now still on the branch. I also switched out the test clips. So, originally, I proposed to use ears, which has this non con con commercial attribute in it. And I switched to VCTK, which has a more has a permissive license. It's at two locations published under two different licenses, both of which I think would be okay. So there is still this fundamental problem with blind bandwidth extension that it aims to generate a plausible high bands and not, so to say, the exact height of bandwidth, which is an ill posed problem. And this is the ultimate reason why it says should and not must. But as I will show later, it's still useful to have that test. And actually, I I found a bug in OPUS one point six point one using that test. So yes, next slide, please. So these are basically the success rates per clip on this test set of 1,000 items selected from VCTK. And, I mean, what you can see is basically that if the quality of the low band is already good, then you actually have you're actually pretty close to you actually saw past ninety percent or ninety five percent pass rate. On the other hand, if the model is very bad, you get fairly low pass rates. I mean, it goes up to sixty four percent, but it's quite a lot smaller than 90. Also, so if you look at six kilobits per second, if you use the original silk outputs, then you get to 20%. But if you apply some enhancement, even just lace, it's already 98. So the proposal is we will say when the quality is good enough, which means for 20 frames, the bitrate is at least nine kilobits per second, then we mandate or we say should pass more than 90%. And if it's ten millisecond, then the quality is lagging a bit. We we say it should be larger than at least 12 kilobits per second, and then the high band test should pass for 90% of the clips. So that kind of separates the good from the bad here. Next slide, please. And now so in the end, I picked one of the failure cases at very high bit rate to kind of demonstrate that point. So left, you have the original. It has not much energy in high band. Then you have this BBWENEet extension, which when you look at it, that isolation is completely plausible. And then on the right is basically the bad model that has the distortion, and that one adds a lot of noise. So the middle one sounds like speech. The right one, yeah, sounds like noisy speech. Mhmm. Okay. Next slide, please. Then there are also two STP parameters for the OPUS media subtype that are suggested. The goal is to let the sender know how much enhancement improves quality. And the second one is how much it would extend wideband speech. So for that, it introduces two new optional FMTP parameters. I see a question.

[01:02:04] Jonathan Lennox: Yeah. I'm trying to understand what exactly an encoder will do with this. Is the idea basically saying if I would normally send 12 kbps and the because of this, I know

[01:02:15] Jan: I can

[01:02:15] Jonathan Lennox: send 12 minus something kbps and and how would I do that calculate? That

[01:02:24] Jan: that that would be the idea. I mean, for real time communication, the encoder might do a lot of optimization depending on the available or estimated bandwidth. Right? And, basically, what you would like to know at the encoder is if I send what what call it quality does the decoder get out when I send this? And, basically, enhancement changes that calculation. K. So, basically, I mean, encoder and decoder are not mandated to do anything based on that. It's informative. But, I mean, you would base you could basically decide to give less bits to audio because, you know, it will be better, give a bit more to video or whatever you like.

[01:03:02] Jonathan Lennox: And then there's no problem. I mean, I'm I I mean, this is I get so I'm just also thinking of how a SFU would handle this. I assume just dropping it is fine. Is there any way to meaningfully so, basically, is the lowest value of this? If you have a fleet of opus decoders and you wanna give the worst case, is it the lowest value of these?

[01:03:28] Jan: That's let I mean, it's up for the implementer. If I was the implementer, then the lowest value would be what's what goes to the sender sides. I mean, what's what's the point of degrading the experience of 10 people if there is one who will get the same quality out as before? Yes. But it's ultimately up to the sender

[01:03:50] Jonathan Lennox: Okay.

[01:03:51] Jan: What he does with that. Okay. Thank you. So there the the names are speech enhancements, and, basically, it's a factor of zero to 40 indicating how, so to say, you pick six kilobits per second as a reference. And then if you enhance that to what bitrates does this correspond, normalized in a way that 10 gives you 12 kilobits per second, which is would be already pretty good. And 40 is just I think it goes up to 30, which should be transparent for silk. And then the second one is extended bandwidth, which gives the bandwidth in kilohertz of the blindly extended signal when you send white band speech in.

[01:04:48] Jonathan Lennox: Yeah. That's not one other pedantic point. These aren't I mean yeah. But these will primarily be used in the SMPF-f-f-dp, but, technically, they're parameters on the media type.

[01:04:59] Jan: Yes. Yes. Thank you. Yeah.

[01:05:05] Moe: Another minor nit. Is that this will eventually go to the STP director and you might get may get some similar comments, but you may wanna think about shorter shorter code points because STP is notoriously bloated, and the the exchanges are getting unwillingly big. So maybe something like EXTBW for extended bandwidth. Maybe this speech ENH or something like that. Just considers the gravity on the on the STP because that all goes over the wire.

[01:05:39] Jan: Those those are already shorter than the ones I proposed earlier. I

[01:05:42] Jean Marc: can Really?

[01:05:42] Jan: I don't have I don't have a strong feeling. I can I can make them even shorter? I mean, if so to say I kind of like when it's you understand what the parameters mean by just reading them. But if, so to say, is of the essence, we can also use abbreviations.

[01:06:01] Moe: I think it's hopeless to read an STP body just by unless you're a member of the SDP directorate, it's pretty hopeless for a developer to read a full SDP body and and understand all of the code points there. It's admirable to make it human legible, but I think you may get you may get comments about, you know, try to pick some, you know, some shorter code points. Okay.

[01:06:30] Jan: Completely fine with that. Yeah. Okay. Then next slide, please. Security considerations, I openly admit these are suggestions by LLMs, but it seems to make sense. So memory safety is kind of clear. Founded complexity, I think, is an interesting one. So less noise and BBWENEet, they basically have the same compute load for every frame. That's if you would add something in there that has different complexity based on the signal quality that would, so to say, expose a risk that a crafted bit stream could spike your CPU. On the other hand, you can shut it off on the decoder side. And, I mean, model integrity is also a no brainer. I mean, OPUS provides a mechanism that you can distribute the model weights independent of the of the actual Opus binary. And in that that case, basically, you have to ensure the integrity of the weights that you pass in. Yeah. Not sure whether that makes sense to have it in or whether that's, so to say, too trivial to even put it in. But that's what it reads at the moment.

[01:07:57] Moe: Yeah. And like we noted on the DRED draft, it probably would be good to update this draft with the final model hashes that you post as artifacts. That way there's some, you know, developer integrity check possible.

[01:08:14] Jan: Oh, I mean, those ones, basically, we won't I mean, the models are not standardized, so we won't officially release them. But, I mean,

[01:08:29] Camille: basically informative.

[01:08:31] Moe: The models

[01:08:32] Jan: because the model are yeah. Are completely informative. Basically, you can use any kind of model that passes these tests. You can use it. And if you build, so to say, a binary for distribution, then you should, so to say, hash the whole thing and also provide the hash. But we can't do it in the RFC.

[01:08:54] Moe: What? We can't do what in the RFC?

[01:08:56] Jan: We we can't provide hashes, I guess, because also it might yeah.

[01:09:04] Moe: Even if they're informative artifacts, it may be useful to just freeze and have a, you know, stable reference for the informative artifacts and have a hash of those in in the spec too so that, you know, a developer that wanted to, you know, test Mhmm. You know, test the enhancement with with that example model, you know, gets a a reasonable result?

[01:09:29] Jan: Okay. Two two things. I mean, first of all, you can I mean, so to say, can build it just from OPUS, and that's downloaded even by hash and hash validated? So that mechanism is there. The second thing is if we publish the artifact, that won't probably won't have the weights, I presume. But I will let comment on that. And the third thing is basically the way OPUS is implemented, it would put all of the different DNNs into one binary blob for delivery. So that would be the one that one could hash, but it's not clear how you would hash the these weights. It it it's except for just hashing the c files that contain the weights.

[01:10:29] Jean Marc: I mean, I I think we could do essentially the same thing as for DRED with the exception that everything in DRED, there's some stuff that is normative, some is informative. In the case of speech enhancement, it would be all informative. But you could have the the enhancement draft have a informative link to the DRED draft about the description of the format and everything. And then it just has the blobs in the it includes the blob in the in this enhancement draft, and we link to the same TAR file for for the source code. That makes sense?

[01:11:26] Jan: It does, but it's a slightly different problem. So you're basically talking about the tarball to include that that has, so to say, all the c files of the model? Yep. That one talks about when you take that power with c files and then you build the, basically, the binary that you will put into your encoder decoder at run time.

[01:11:55] Jean Marc: Yeah. The blob file, we could also separately just like for DRED. Like, in the in the DRED draft, the source code would still have all the weights, but the the blob file would contain the

[01:12:08] Jan: It would also be so it also contains the blob file, you're saying?

[01:12:14] Jean Marc: There's, like, a source star ball and it contains all the weights, but there's also a blob file that has just the weights in them.

[01:12:22] Jan: Okay. And that blob file could I mean, that blob file probably contains these models even. Right?

[01:12:28] Jean Marc: The blob file has the models. Yes. I mean, it has the all the weight. It doesn't describe the models and so on, but it has all the weights.

[01:12:36] Jan: Yes. Okay. Then we can basically put this into into quote blob and kind of add add these hashes. That's that's possible. Okay. Thanks.

[01:12:48] Moe: Good. So it sounds like we're gonna try to mimic the same thing that that happens in the DRED draft here for speech enhancement.

[01:12:54] Jan: Yes. Okay. Good. Okay. I mean, basically, then the next steps, collect and incorporate feedback, we already added. If there is more, please share. Then so my my goal is basically to incorporate this until the next meeting. I it would also be great if someone could cross check the evaluation, which just means running the script. If they have volunteers for that, I still need to basically land the latest changes on main. That's now at the ose_ietf120 branch. And then the test vectors would have to be uploaded somewhere, and they are a bit chunky. It's four gigabytes, but it's less than training data. Much less than training data. Is that gonna work?

[01:14:03] Moe: We hope so. I don't remember the exact cutoff that I don't know if I don't think I gave us an exact cutoff, but I I suspect that four gig will be hostable.

[01:14:16] Jan: Okay. That's fantastic.

[01:14:18] Moe: That's it.

[01:14:25] Jan: Okay. Yeah. Thank you.

[01:14:27] Moe: Thank you. Alright. And so now Camille, are you on? Hello. Hello. Can you hear me? Yes. We can hear you fine.

[01:14:49] Camille: K. Let's see if we can get this video going. Okay. So this is an update for the test battery for DRED. Next slide. We've adopted that individual draft that Laura and I published into our working group documents. No changes there other than just some editorial fixes. Not planning to make further changes to the evaluation battery. And the utility of the battery is for supporting reproducibility and benchmarking of the standalone DRED. And that same battery was used in the new product challenge last year, and some further details were published at ICASSP this year. And these results, we are in the draft and we reported previously, so this is just for reference. Nothing new.

[01:15:54] Moe: Next slide. Just real quick. Is that DRED version, is that the latest gen mark that's the last final version? 1Point52, is that the final version?

[01:16:08] Jean Marc: Sorry. If what is 152? No. It's candidate b that is the latest version that we have.

[01:16:19] Moe: Oh, candidate b. Okay.

[01:16:20] Jean Marc: Was the pre was the

[01:16:22] Moe: Oh, I'm sorry. Okay. Was just looking backwards. Okay. So the bottom is the is the most recent. Okay. Yep.

[01:16:28] Jean Marc: Candidate b and candidate a, we did not select, based on the results, basically.

[01:16:32] Moe: Candidate b is the current draft's final weights. Yep. Great. Thanks.

[01:16:40] Camille: Alright. Next slide. So for future directions, we are looking for input from the working group. I guess on the first question, if there is no changes to Dread that need testing, then there's no need to rerun the battery. But there there is another possibility. We are wondering whether the working group would see value in in the in us investigating some objective metrics and recommending which ones and under which conditions have highest correlation with the test battery for DRED specifically and then reporting reference numbers for such metrics so that anybody reproducing and benchmarking in the future doesn't have to go all the way to test battery, but can run readily available metrics and compare the numbers. And we've we've we've done some work related work in this direction recently on neuro audio codecs, evaluating a dozen or so neuro audio codecs over 45 different reference and non reference intrusive and nonintrusive metrics in clean condition specifically that's been published. And then we are working also similar work in noise and reverb that's that will be published later. And the conclusions there might hold for Dread or they might not. Jean Marc mentioned in the past, DRED is not, by design, not generative in nature. So it might it behaves differently than some of the other knacks, so we might find that there are some differences in conclusions or recommendations. But either way,

[01:18:47] Timothy B. Terriberry: we

[01:18:47] Camille: could give guidance on this and and produce those results if this if their working group sees this as valuable.

[01:19:02] Moe: So on on the first item, I assume that if there are no changes from candidate b until ISG review, then I don't think we'll need to rerun the test battery. I don't expect ISG to ever have a big enough substantive issue to cause the DRED model to rev its weights. So I think you're probably safe there and not having to rerun that. But but I guess we'll have to keep this open until we actually pass ISG review. On the objective metrics, that sounds very interesting. But I'd like to hear from the work group if if anybody would attempt to run those objective metrics if they had them available, or would we even try to incorporate them as part of some, you know, CICD, you know, continuous kind of validation of of Opus changes? Go ahead, Jean Marc.

[01:20:07] Jean Marc: I'm not sure I fully understand the idea of these objective metrics in the sense that sort of the the subjective tests you've done already show that DRED works. And to test compliance, there's test vectors on DRED. So I'm not sure what the objective metrics would you can you?

[01:20:30] Camille: Yeah. So so compliance wise, I guess, the test vectors suffice. But, you know, objective metric would function in terms of being able to predict subjective impressions. So if if somebody was improving DRED or or reproducing DRED and it had some artifacts, they could gauge maybe without running the listening tests, which takes some resources to integrate with crowd source platform and is it costs money. They could just readily take an available metric and get a prediction, and we could quantify how well correlated the metrics are. Although, you know, it's it's always needs to be taken with a grain of salt because depending on types of distortions and and and the inputs, the the correlation might not hold. But, I I guess, the input, would keep the same. We would give a reference set. But still, the the types of differences that somebody working on an algorithm like that might if they weren't included in the correlation evaluation, they might still differ. But yeah.

[01:22:07] Jean Marc: I mean, could that would be for testing an encoder or, like, what would what would these tests actually validate?

[01:22:18] Camille: Uh-huh.

[01:22:19] Jean Marc: Just trying to

[01:22:21] Camille: I think it's more the quality output of the decors or the fire gun. And it's it it would be quality. So so, like, the battery has quality in clean and then degradations in mild noise and reverb. And I guess we didn't we haven't extended this work to intelligibility. So the only thing the the metrics would offer is predictions of quality and intelligibility that are hopefully as well correlated with the subjective test as possible from a pool of metrics. So we would essentially say these metrics in clean conditions at high quality are the ones that gonna give you the highest correlation. And these metrics at poorer input quality might be more suitable, and it will they those will give you decent correlation with the outcomes of the battery. But that that's all they would offer. And then but, you know, I I see your point that the battery covers the subjective prediction and then test vectors cover, like, validation. So yeah.

[01:23:48] Jean Marc: No. I mean, I guess it so there's many there's many steps in the DRED process, and the very first one, like, the feature decoding that is very tightly specified normatively. So I don't think your objective metrics would be useful there, but the Vocoder is much more loose and the actual integration is even looser. So I guess if you wanna test end to end, then this could be useful for testing and implementation, I guess.

[01:24:19] Camille: I mean, it it would be aimed for for the whole stand alone codec. And I guess since the decoder people could create their own, I don't know. Higher capacity, lower cap whatever. As long as the you know, they're using the same features, they could tailor the color. So then without having to run the subjective test battery, they could get an indication of quality and compare with the baseline model, essentially. Okay. With the baseline end to end model.

[01:24:53] Jean Marc: In that case, it would probably be useful to, at the very least, compare with the the there's a dred_compare tool that is also meant to be able to do end to do, like, the end to end evaluation, and it's probably not great. So it would probably be good to include that in your objective metrics.

[01:25:19] Camille: And, sorry, Jean Marc, what does it what does it do, the tool?

[01:25:26] Jean Marc: So it's called dred_compare. It's used for the DRED test vectors. Uh-huh. It's doing a bunch of different things, one of which is to validate the features, which is irrelevant in your case. But one of the things it does is take audio from, like, an end to end system. Like, you decode and it compares to the encoded like, we you decode with loss and everything, and it compares with the original. And it's there as a sort of sanity check type of thing. So it's probably useful if you include it in your objective metrics evaluation, if only to show that all of these other metrics are better than the dred_compare tool for evaluating end

[01:26:07] Moe: to end. Just to clarify for, Kyle, because I'm not sure he was here in the very early presentations, on this. But the dred_compare is basically to make sure that DRED doesn't turn a yes to a no or something. Right? It's not to actually compare quality. It's not to compare that the quality of the speech is good. It's more like that it it's close enough to the to the source. It didn't perturb the source in a very

[01:26:33] Jean Marc: three both of them. There's three tests there's three tests in the test vectors in the DRED test vectors. The the first one just checks the feature decoding. That one is tight. The second one tests that the Vocoder interprets the features correctly, like, that the spectral envelope is the same as intended and that sort of thing. And the last one had bit streams with faults in them that gets filled with DRED, and then dred_compare will compare the audio to the original. So and that one, you know, I made passing a should, and, you know, this tool is, like, very approximate, but I think it would be useful to include it if only to show that, you know, these objective metrics are better than the DRED one. And if any of the standard objective metrics are worse than the DRED one, then they're more than useless.

[01:27:31] Moe: The do the does the dred_compare actually generate a metric, or is it a binary pass fail?

[01:27:37] Jean Marc: No. No. No. It actually generates a metric.

[01:27:41] Moe: That is roughly comparable to quality?

[01:27:44] Jean Marc: Very roughly, but yes.

[01:27:49] Moe: Okay. So I guess the relative question would be, are these other are these new objective metrics closer to subjective and more aligned with quality, you know, than than the rough quality that you get from from the from the dred_compare? I guess that's the key point to whether or not that's useful for for future direction in the test draft.

[01:28:15] Jean Marc: Yeah. That's the that's the question, and I would assume that the answer would be yes for most, if not all of them. Because dred_compare is a pretty crude tool that is meant to do, like, this one thing.

[01:28:27] Moe: Yeah.

[01:28:29] Jean Marc: You're not trying to boil the ocean Just enclosure, including it sorry?

[01:28:33] Moe: You weren't trying to boil the ocean and and solve a general objective quality, you know, objective speech quality metric. Right? You were just trying to validate DRED. DRED doesn't distort signals too much. Yeah. DRED is faithful, but not the best possible quality.

[01:28:49] Jean Marc: Yeah. So if if anything, like, if you show this objective thing is, like, much more correlated than the dred_compare, then it's, you know, it gives some useful feedback about, you know, use this if you want better evaluation and that that sort of thing. Mhmm. I think it's sort of the bay the lower baseline or something that when it comes to DRED testing anyway.

[01:29:17] Moe: So, Jan, were you waiting to say something? You're muted again. You're you're meet mute echo muted this time. Okay.

[01:29:25] Jan: Yeah. Yeah. Okay. So the good kind of mute. Yeah. So I definitely find it interesting to make recommendations about or report, so to say, correlation of objective metrics with quality. I would be a bit cautious about saying this is the value that you should expect to get. I mean, DRED is incredibly I mean, first of all, so to say, DRED involves things like packet loss and all these things. And my experience is a bit of a warning. Right? The one clipping changed in the decoder, and suddenly things look different. That might show up here as well. And then, so to say, having a useful test for engineers who would integrate something like this, it will be kind of difficult because you can only evaluate the Dread performance if you have packet loss. And then if you really have an end to end integration, you wouldn't even have aligned an aligned reference anymore because the whole thing would go through that eq, and you would have gaps in your audio when, so to say, the offer expands. You would have accelerate in there. So if we can say something like I don't know. I think is that the study that where WARP-Q came up on top, out on top?

[01:30:48] Camille: Yes. Yes. That's the one. Yeah.

[01:30:49] Jan: That that's actually interesting because in the early days, I used WARP-Q, the predecessor, I think, to that one in a hacked way as a reference metric, and that worked also best for, so to say, ranking opus, lace, and no lace. So if we can test this on in a relat realistic scenario and say that correlates best, then I would find this very valuable. But whatever values we put there, we should really say you should take it with a grain of salt.

[01:31:20] Camille: Not a grain of salt. Okay. Yeah. And so, Yan, do do you because I was thinking initially to to keep it simple and repeat what we did before, meaning standalone DRED evaluation. But you are talking about end to end with packet loss and all of that being divided by the party.

[01:31:46] Jan: Yes. I I mean, in a certain way, you want if you want to really know whether your implementation works, then you have to Yeah. Observe it while it does what DRED is supposed to do.

[01:31:57] Camille: Right? Exactly.

[01:31:59] Jan: I I mean, there is an open source patch to WebRTC that basically does the integration that had also been used for these simulations that are on the open release page. And, I mean, one probably should take that data and then run the objective metrics on there because that that's that's most realistic then. I mean, just just so to say validating that thread inside of the codec works, that kind of seems like the limited test case. But when it's if it's guidance, so to say, for people who really want to integrate it, then they would only like to test the full system.

[01:32:48] Camille: So may maybe we can take it in stages and design it. We need to and start with simple and then progress further. Okay. I'm just trying to think whether we would be leveraging anything that we've done already in in the stand alone case, versus in the new case, we would have to because in the stand alone case, all we have to do is really just run objective metrics. We already have the subjective test results, and we can check. In the other case, it is a considerable undertaking to both well, get the tools that produce the realistic scenarios and apply the codec for what is was designed for catering to packet loss, etcetera, then run the subjective tests and under objective metrics then to the correlation. So, yeah, it it's considerably more work. So so possibly in the future

[01:34:01] Jean Marc: If I may suggest something, like, I think including something like NetEQ sort of blows up the problem because then you start playing stuff faster. You'd have zero alignment. For objective metrics, that would be, like, a huge a huge undertaking. But I think, like, OPUS demo level end to end where you still have losses, but at least everything remains synchronized, like, I think that may be a decent compromise. Like, you you should be able to do better than what dred_compare does in terms of evaluating an implementation. You would still evaluate whether the encoder is doing a good job, whether the offsets are computed correctly, all that sort of thing at the opus level. I think that makes sense, and it's kind of testable, not completely you don't have this alignment problem or anything.

[01:35:03] Moe: So I

[01:35:04] Jonathan Lennox: think that

[01:35:04] Jean Marc: would be a compromise.

[01:35:06] Moe: Yeah. I'm sorry. This is a great conversation. Maybe we can continue it over the list or on Zulip on on chat. We're a few minutes over, so I just wanna try to close off this final position. What I what I heard is this is a work group draft now. What I roughly heard was we don't have work group consensus on on what the objective metrics are and should be and whether that would be useful or not. So I think so I think, Amil, I think the feedback to you would be as as editor of the draft, do not add anything to it yet. But if you wanna come back later and bring forward some some proposals, I think the group is is, you know, open to hearing proposals on objective metrics, but there's just not consensus about what to adopt or, you know, or or rev in the draft yet. So maybe just propose something in next time, and we'll see whether or not it's something that may be useful to add. Sounds good. Alright. Thanks, everyone, for the time today. And we will start a couple of working group last calls on the list. So be looking for those to to give your input on them. Thanks everyone. Have a great weekend.

[01:36:22] Jean Marc: Thanks everyone.

[01:36:23] Jan: Thank you as well.

[01:36:24] Camille: Thank you.

[01:36:25] Moe: Bye. After Mark.

[01:36:50] Jean Marc: So

[01:36:54] Moe: let me ask you as someone who I value as a very reasonable and astute evaluator of things. What do you think about doing the top tracks in mock? The top the top end tracks filter. I hear what the audience's view of of that is.

[01:37:21] Jonathan Lennox: I mean, it seems like it is a leaker to me. I I am a bit of