**Session Date/Time:** 21 Jul 2026 14:30 [00:00:05] **Sarah**: Yeah. I can you back that out so I can get control again? [00:00:15] **Johnny Ryan**: Oh, sorry. [00:00:17] **Sarah**: Yeah. If you can [00:00:18] **Johnny Ryan**: release that. Sure. [00:00:27] **Sarah**: I can't share slides from here. Oh, my the echo has frozen. K. Let me rejoin. Okay. Are we ready? Yeah. Okay. Welcome, everybody. This is the PEARG meeting of the IRTF. This is the note well. If you are in the room, you are bound by the note well. We assume you have read it. If you have not, please go and do so now. If you are in the room, please scan in using the barcode with the light client, the on-site client. We use the numbers of people who scan in to plan room sizes for future meetings, so it is very important that people do that. And if you're remote, please keep your audio and video off unless you are asked to send it. Our agenda for today, we have a minute taker at Laura in the front row. Thank you very much. Our agenda has one late minute addition, which is an update on a draft that was proposed That was proposed to me. We are very excited to have two excellent main presentations today. The first one is on real time bidding, how advertising interacts with your privacy by Johnny Ryan. And the second is on large scale online de anonymization with LLMs by Simon Lerman. And, after that time permitting, we have a short presentation on source privacy after ECH. We wanted to quickly give an update to the group on what's happened this week regarding the draft that was posted to the mailing list, I think it was last week, around a proposal for a ciphertext inference tool for MASQUE. We were initially a little bit unsure about where this fits into our charter, and so the recommendation was for it to go to SecDispatch. The authors who are in the room over here at the front did a presentation, and the conclusion there was that the content of the current draft could potentially be split into two, and that the one part that would live in PERG would be a privacy analysis and an evaluation of the technique, in particular, looking at the applicability of FHE as proposed in that draft. So, we've talked with the authors and we're going they're going to update the document to redefine the context for it so it does fit within our charter. So, we look forward to that. And then, the second separate part of that draft that actually deals with the pro to call integration will probably be passed off to a discussion with the MASQUE foundation. So, we look forward to having that work coming through the group. Before we started our presentations today, we wanted just to take one moment to frame the discussions that we have because both are really quite sobering presentations on the current state of the landscape. We wanted to talk about potential research topics and a community agenda that we could have either here in PEARG or in the wider IRTF. And so the kind of topics that we were thinking about and we're hoping that people can have in mind when they're listening to these presentations is is moving forward, what we could could we do around concrete measurements and analysis in this arena, the threats to anonymity from hyperscalers and AI? Also, could we do a more formal threat analysis or comprehensive threat modeling? And then following that, what can we think about as potential mitigations, be they technical or sociotechnical in this context? We're interested to know if there is appetite for a workshop on this topic and considerations for research in this group and beyond. So, we we hope this will generate some discussion and a new avenue for research for the group. So, with that, I will stop sharing my slides and I will hand over to our first speaker, Johnny Ryan from In Force. Please go ahead and introduce yourself, and hopefully, your screen will share nicely. [00:06:12] **Johnny Ryan**: Thank you very much, Sarah. Yes. I'm in the process of trying to share my screen. Okay. So in the theme of doom and gloom, you should see a big black screen in front of you. Sivan, thank you for inviting me to this, and, Sarah, thank you for that. So what I'm going to show you is a deck that is not new. I have had versions of this deck for eight years, maybe longer than eight, depending on the version. So this is not a new problem, but I do think this problem is probably still how most content is paid for online. So I'm going to explain the basics of how real time bidding works. And maybe the best way to do it is if we play a kind of a hypothetical game where if everyone in the room imagines we are working for a company, and we want to sell a product to a particular person. So let's imagine we want to sell some sort of dieting pill, right, to a particular person where I live in Dublin. And we're looking for someone with a high disposable income, but they have clearly a diet that is full of high calories. That's the type of person we're looking for. Now we are the marketers. We're represented by that dollar sign on the left of your screen. We're trying to get to the visitor to a web page or user of an app on the right hand side. So here's the process, and I'm gonna take you through it step by step. So let's imagine that this pill is called, I don't know, health pill, and we already have some customers. Maybe we have 10,000 customers in this country, in Ireland. So we will take those data to a company called a demand sorry, a data management platform, And they will store the data for us, and very often, they will offer us the opportunity to augment the data we have on our customers with even more data about them. So what happens there is that the DMP will share the information they have about our customers with various data brokers, and they'll say, can you match any of these people and supplement our understanding of them? Do they have high disposable income? Are they eating fatty foods? That kind of thing. Whether or not we go through that step, we now have some data in our DMP, and that's useful for us because the next thing we need to do is to equip our demand side platform. There will be today millions of auctions even in Dublin, the city where I am, for advertising spaces on websites and apps, and we clearly cannot be present at those auctions to make bids. So instead, we're going to have an agent who's present, and that agent is going to behave in an automated way. But they need to know who they are going to bid for and how much we're willing to spend. So when a person hits a website, how much are we willing to spend to get R ad in front of them? And we may even put a different price on different people. If this person is already in our DMP, there's someone we've seen before, maybe we're willing to spend more on them, or maybe we'll spend more on lookalike people who are very similar to the people who we already have data on. Okay. So we tell the DSP or rather our DMP tells our DSP what it needs to know, and now we are ready for someone to visit a web page. I'm going to leave apps out of it and just talk about websites, but it's fairly similar. So along comes Morris. I've invented a person called Morris. Morris loads the Irish Times website. That's a a newspaper that I once worked for. And in the space of milliseconds, the editorial content is is sent to his browser, and there's an empty rectangle because The Irish Times has not yet sold in a direct deal. It has not yet sold the advertising space for Morris to I've just noticed, by the way, that I've set my timer for fifteen hours and forty minutes. So I'm going to pause this presentation, and I'm gonna change that timer. And this might be the best thing that I could do for anybody. Okay. Timer. Where are you? Okay. Forgive me. I now have even worse news, which is that I cannot find my timer. So if I start running out of time, you know, please make distressed sounds in the room. Okay. So Morris is loading the editorial content of this article he's reading. And The Irish Times has a sales team who phone up people and do advertising deals with them. But very often, it won't have sold all of its space, and most publishers and most apps don't have a sales team. So in those cases, in the majority of cases, that rectangle at the top of Morris's page, we're just picking one ad, is empty. And so in that moment that the ad is loading, in the space of about two hundred milliseconds, here's what happens. Information about Morris is going to be sent to a supply side platform. That's the counterparty to the demand side platform. Now I've euphemistically called this cookie, but we'll get into what the information is. The supply side platform will run an auction itself, but it will also solicit bids from other similar auction houses. So you end up with an auction of auctions where one or many ad exchanges are receiving information about Morris, where he is, what he's looking at. And that will be sent by each auction to all of the DSPs, the demand side platforms that are sitting in each of these auctions. I'm going to show this to you in another way shortly. So if that sounded a bit convoluted, it's it's okay. Now tens or hundreds of different DSPs have received information about Morris, but our DSP knows that he's the perfect person for us. So we bid more than everyone else, and our ad appears in that top rectangle on Morris's page. And you'll even have noticed many times loading a website. Sometimes the ad jumps in and it pushes the content down the page. So for anyone who's not blocking everything, this is what the normal process is for RTB. In general, and this is a bit of a euphemism, there is a opportunity for many of the parties, I'm simplifying this, to synchronize their identifiers for Morris and thereby synchronize what they're now learning about them with what they already know about them. So let me go through that two more times so that it's super clear. I arrive on the dailybugle.com. The daily bugle has an ad slot, and so it throws me to the wolves. It sends information about me to, in this case, four ad exchanges, and each ad exchange sends that information on to four or five DSPs. And the theory is that the right DSP with the right advertiser will pick the right ad and bid the right amount to show me the right ad at the right time. So this is a very compelling offering for advertisers or at least it used to be ten or fifteen years ago. Now here's the main privacy problem. We don't know where the data go. Once they leave the page, there's no guarantee at all that there's any control over them. And we can know roughly what does happen because there's one of these DSPs that was the subject of an investigation by the CANIL, the French data protection authority. And this company was a victory. Now I used to work in this industry, and I had never heard of these people. They were and are, I think, still tiny. So 3,500,000 turnover in the preceding year. And yet when Kaneil investigated them, they found that they had hoovered up 67,700,000 people's data just by sitting in on these auctions. The last time I visited Vectary, their website, I I found this on their website. Privacy is hard coded in Vectary's DNA. And this is the normal industry spiel. Clearly, it's not actually the truth. If you scroll down this page, you'll see two interesting claims, though. Vectary says we only store 30%, and we dump everything after a year. So that begs the question. This is back now in in 2018 when they were investigated. Were they receiving a quarter of a billion people's data just by sitting in on ad exchanges? This teeny tiny company. And as we'll see when we get into the question of scale, that's entirely possible. So I'm going to death by PowerPoint you now, and we're gonna spend thirty seconds doing something very courageous. We're going to load a website. So you arrive on a web page. The website's ad server connects to the SSP, which sends information out about you to at least one ad exchange, which broadcasts each ad exchange broadcasts that information to tens or hundreds of other companies called DSPs. There's generally an opportunity for some or all of them to do some form of a sync, And then there's a whole lot of other plumbing that you don't need to worry about for the ad to be shown. Now the purpose of this diagram is to make you think that's an awful lot of arrows, and it is. So the question is what's in those yellow arrows? What's in a bid request? And we can know, actually, not by instrumenting a client because often a lot of this is behind the scenes server to server. Some of it happens on the client. But we can know because, actually, up until recently, there were only two industry standards. Now, actually, there's only one because Google also uses the IAB standard. So the IAB, you will be familiar with. This is one version of the open RTB standard or protocol. So there are several of these, but they're all largely interchangeable for privacy purposes. If you scroll down this particular one, it's about 14,000 words. But if you scroll down to the very bottom, you'll get what is described as a summary example of what could be in a bid request. Now there's about 600 different variables or fields that can be in a bid request, and it could be something as benign as how wide is the rectangle, how high is it, can it display video. Not really something that should trouble us too much, but there are plenty of other things. Often, there's the entire URL, you know, embarrassingwebsite.com slash excruciatingly socially awkward article. So, So, generally, you get the entire URL, but not always. You've got two different IDs. There's the ID from the ad tech firm and the buyer user ID, that second one in this example. That is if our DMP through our DSP had preannounced to this to the SSP that we're looking for this person. Do you have him? Here we have this person's age and gender and various other stuff. Now this this document this standard, I think, dates back to 2018 or 2019. I can't remember. And so we're getting things like the user agent string, and we're getting the lat long. Google and various other ad exchanges may truncate some of these things. I know that Google, for example, truncates the IP address and some of the other stuff. But this isn't the entirety of what can be sent. So in general, this is what's sent, and I just want to draw your attention to another related standard. This is the IAB content taxonomy, and this can be sent these these codes can be sent in a bid request too. Now what I'm showing you is a page that I printed out years ago. I was ticking particularly egregious fields on the right hand side, and I found myself switching to circles with exclamation marks when I encountered IAB seven hyphen two eight incest slash abuse support followed by incontinence, infertility, and so on. Now IAB now claims that this particular set of classifications of what a person's reading, which can go out in the bid request, that this is depreciated. But I found these particular classifications as segments that were purchasable, you know, through through the RTB system and through data brokers. So they are out in the wild, and they they are still used. But the supposedly cleaner version isn't particularly good either. Okay. So let's let's just take stock on what's going on here in one summary slide. It is a MERV. It's a massive broadcast of data where tens or hundreds of DSPs are receiving maybe some very spectacularly sensitive information about people. So let's think about scale. If we take one single ad exchange, it's it's hard to know how many could conceivably how many other parties could receive data from a single ad. So let's just zero in on one ad exchange and try and get a sense of scale. So I'm gonna take a look at Google's public documentation. This is for their business in The United States. This is the list. It's live right now of the companies that it can send data to from its its RTB exchange. And I'm gonna keep talking until I get past the letter a, and that is taking me longer than I remember. I'm gonna stop talking. We're on a d. So this is a broadcast in the old sense of that of that word. We think of broadcast for radio, but, actually, you know, this is an ancient agriculture. It's it's where you put your hand into a bucket of seed and you scatter it on the wind in the hope that it will land on a fertile patch, and that is what real time bidding is. Now briefly, let's take a look at the law because sometimes it actually brings clarity. The GDPR has been called by some incredibly complicated actually, it's really simple in some ways. It's it says very, very clearly, you're not allowed to have a data breach. General data protection regulation. Just protect the data. Now industry documentation, and this is publicly available online right now if you do a search for the document, says thousands, which I think is maybe an overestimation. Thousands of companies can receive the data from one of these bid requests, one of these broadcasts, and there's no way to limit what happens to the data once the horse has bolted. For each ad, once that broadcast happens, we don't know what happens next. Now I haven't updated this in a while, but I think it's probably the same. This is Google's terms. If you are going to sign up to have a seat on Google's auction, you have to agree to a a document, this this program guidelines document. And they refer to companies that are sitting on their auction as buyers, authorized buyers. So you as a buyer, you are not supposed to use callout data. That's a bid request, real time bidding broadcast data to do this other stuff. And if you intend to break this rule, you know, you should fax Google at your earliest convenience. So this is self regulation's inbred child with another piece of self regulation. This is nonsense. And in the run up to the GDPR, when I was still in the industry, I remember there was a big focus on this idea, this notion of consent. But at the time, the IAB had written to the commission, the European Commission, and had begged for a loophole in the e privacy regulation. And they included this document where they say, we don't really know what happens to the data. Or, actually, they they don't they don't say that. It's impossible for the person visiting the website to know what's happening to the data is what they're saying here. But, obviously, it's impossible for anybody to know because there's no control. And when you don't know what you're going to do with data, you actually can't ask someone for consent for that because it's a data breach. So, unfortunately, we ended up with this catastrophic mess called the transparency and consent framework. And I have been litigating to try and kill that for quite some years. So we have consent spam, which is a layer of kind of compliance theater. Right? A source of quasi legality that sits on top of this data breach. But, actually, it's a data breach, pure and simple. It can't be consented to. This all this does is antagonize people. So back to that question of scale. A few years ago, I was able to get my hands on data about how many advertising opportunities there were using Google's ad buying system for real time bidding, which includes not just Google, but other auctions too. Now it doesn't include Meta has its own system and Amazon too. It doesn't include them. But what you end up with is astonishing figures that, on average, at least at that time, A person in Arizona would be exposed in this way 782 times a day. A person in Florida, 870 times a day. These are staggering numbers. I I suppose they do come down because the same person may be subject to multiple broadcasts on a single web page of the same website because there are multiple ads. But still, we're talking about an absolutely enormous data breach, and it's happening every day. It's not just the biggest that I can think of, but it reoccurs. And the upshot is that everyone then is exposed. And so what do the dossiers that result look like? Well, we can know in quite a simple way. The industry has been quite open in its public documentation about what those what those fields can look like. This is the audience taxonomy version one. It dates from May 2018, which hilariously is the month that the GDPR became enforceable. It just hasn't been enforced at all. That's that's one of the consequences we're living with. So let me take you through through a few examples of what's in this document. Previously, I showed you the content taxonomy, which was about what you are looking at on the website or on the app. But now I'm showing you the audience taxonomy. And this is, I think, a kind of a Rosetta Stone by which multiple data brokers can easily combine their understanding of a person by having common codes for different fields. So the audience taxonomy has, you know, around about 2,000 canonical fields. This isn't all that granular. But those fields include your religion, which is, of course, very sensitive, whether you've a mental health symptom, infertility, STD. Are you worried about losing weight? This would be very useful in my example of trying to sell you that weight loss drug. And then the taxonomy gets into some of the most intimate things about a person, which isn't just about them, but it's about their child. Do you have a child who has special needs? If you do, then the IAB says the code for you is 357, and it's okay to broadcast that code to hundreds, if not thousands of companies. What are your political views? And it is it does astonish me that we are dealing with people who can't spell. What's your income? Now in the case of income I'll just stop there. In the case of income, I am not suggesting that all of these data could come from real time bidding, but they can be applied in real time bidding and combined in real time bidding. I don't know how you would have granular salary levels just from web page visits. Maybe there's a way, but I haven't thought of it yet. Anyway, let's assume we know your income and your health issues. Maybe it's also a national security issue if we know that you work in aerospace and defense, especially if we know that you have a personal debt problem if you're facing bankruptcy. And it goes on. So these are very compromising things. I'm just going to skip, you know, TILA and online gambling, bail bonds, payday loans, and so on. So that's that's the basics of what's available. Now some time ago, I was able to obtain, by posing as a data purchaser, segment lists that we would use as advertisers if we were trying to buy ads for our product, going back to the beginning of my presentation. We would be looking for segments of people who are who have high net high disposable income. They're interested in frozen foods, fast food, high calorie food, and they're interested in weight loss. So we've been looking for segments of those people. Well, posing as a data broker, I was able to obtain lists of segments that were available from different companies. Now I'm going to zoom in on this slide. I understand you can't read it, but I want to explain what you're seeing first. So ignore the first two columns, which are just showing the row numbers. After that, we get the name of the segment. I'm gonna zoom in on that in a moment. After that, we get the the company that that segment has been provided by, and then we get details, a description of that particular segment, and then we get various numeric codes. The numeric codes, I want you to remember them because they're important. Those are the segment codes that you would go to each SSP with and say, I'm looking for this segment. Could be Google. It could be the trade desk. There's several of them. So let's zoom in on one of these segments. This was a particular list from The United States. Actually, I have them for countries all around the world. And you can see that this list from Ad Astra is about or this segment is about decision makers at government organizations primarily engaged in national security. So let me just group several of these interesting segments together from The US in this case. So you're seeing judges, government decision makers, active military, government seniority, chief level officer, CXO, and so on and so forth. Now the segment lists are many hundreds of thousand long, so there's all sorts of personal characteristics and all sorts of player types. The segment ID codes at the top, you see TubeMogul, TURNID, AppNexus, which is a company that was bought by AT and T, renamed Zander, and then bought by Microsoft. And then you'll see the double click bid manager ID, that's Google, and so on, the trade desk, and so on and so forth. So the data are very compromising even if you weren't concerned with personal privacy, even if you only were concerned with national security, you would be worried. And as you go through Google's list, you'll notice if you're in The US, for example, PADFA is a law that says you cannot sell data or send data to so called foreign adversaries. The name Beijing is in many company names on this list. I worked on a piece with Wired magazine that examined whether Google was in fact sending per data in these bid requests to companies in China, and the answer was yes. Now that that thing we were scrolling through showing all the companies that Google might send your data to, if you print it out, it's a 150 pages long. There are 2,051 companies on that the last time I checked. So, again, with the national security lens on, which which shows just how acute the personal issue is for individuals, you can start to see what kind of data are available about a person. From the segments that I have seen, you could definitely find out what a person does no matter how sensitive their role. RTB is often sending location data So you can see where a person is spending their time and where they're traveling. You can see where they live. You can also see their most frequent driving routes. The reason I use that term is that we found a Israeli surveillance company that said it it relied on data from real time bidding, and it offered something that looked useful for a drone strike where your target's most frequent driving route was and when they met with other targets. All of these tags appear in the segment list, personal debt, bankruptcy, gambling, etcetera. But even in some of them, there were German lists where you had Cambridge Analytica type psychographic analysis of people. So were they dutiful? Do they have a low affinity to that? Which means, I think, could this person be open to a bribe? And things about their their life stage. Are they menopausal? Do they have panic attacks? And, of course, you can find all the other basics that you would expect as well. Now the reason why I mention Tinder and Uber here is there's a a remarkable case of a Catholic priest, a Monsignor in The United States. And this individual was outed as gay because a conservative group had spent a lot of money to buy data about priests and their locations. They wanted to see if the priests were staying celibate. This particular priest, the data they obtained about him, I think it was about fifty six weeks of his life, you could see his gay hookups all over the place and the bathhouse he was in in Vegas and so on. And they leaked this to a journal, a Catholic journal called the pillar. And this particular priest was a monsignor who used to often appear on television in The US, and he ended up leaving the clergy. He's currently suing I think the company was PubMatic, which was the RTBSSP, if I recall correctly, for Grindr, and that was the app he was using. So you can tell people states, but you can tell a lot about their lives. So let me briefly and the reason I I flash up next page is that you can definitely tell a lot more too. So I'm going to stop now because there might be some discussion if you're interested in it. But before we do, I wanna talk about a solution. I'm not suggesting that we should not have online advertising. I have worked for a news publisher, and I understand and I value the idea that advertising can sustain newsrooms and can pay for entertainment. And I'm also not against the idea that you could even have auctions for advertising space. The solution, I think, is actually very clear. The whole industry relies on a single specification document, which the IAB publishes, and it defines what fields can and should be in a bid request, the 600 fields or so. And what needs to happen urgently is that that specification needs to be fixed. It is okay to know that a person and you're praying they're a person and not a bot. One of the things this tracking enables is for bots to better masquerade as people. But it is okay to send out information saying, whoever it is that is visiting this page or using this app at this moment, whoever they are, they seem to be in this general area. You might even say, this is the article that this person is looking at. But you've no way of knowing who the person is, and you can never connect what you're learning about them from this broadcast to the next broadcast. There's no way of tying the data together. And that clearly means truncating many of the fields. Last time I checked, I think it was about 8% of them. Certainly no unique identifiers, but it probably also means even slightly slightly distorting, you know, to the tune of milliseconds, the time that the that the bid request is sent to each DSP so that you don't have, you know, a reliable high resolution time stamp, if that's possible for them today. So they they've no way of joining what they know about the same person. But they are understanding that someone in a fairly lucrative area is reading about golf, in which case, I don't know, show them a financial product for for someone with a fair bit of cash. This is the way advertising has always worked, and I think it can again. But this particular problem, I think, involves reforming RTB by changing the specification. So the last thing I'll say is we have a class action right now, and it's against Microsoft in Ireland where Microsoft has its European headquarters, if we succeed in the class action, then that will force Microsoft to amend its its bid request specification and truncate or remove those identifying fields. So it's an entropy question. And if we are successful, that will then apply across the European economic area and hopefully will be a model for the rest of the industry. The address I've put up on your screen is a it's a kind of a basket where I put whatever reports or evidence we have produced over the last few years and an update on our current work and exposes on this and where the court cases are. So on that, I'm gonna stop sharing, and maybe we can have a discussion. [00:40:24] **Sarah**: Thank you very much indeed for that presentation. We have a very full queue. First, Daniel. Please go ahead. [00:40:31] **Daniel**: Daniel, good. Hello. Thank you for your talk. Two part question. Are any of the IAB classifiers or categories exposed to the browser? So could I, as an end user, read out as what kind of person I'm classified? Or, like, follow-up, could I or someone post as a data broker, set up a test page, and thus get our hands on, like, the classifiers, and then expose that information back to the user? Because I'm, like, really interested in what, like, the ad industry thinks, what kind of person I am. [00:41:07] **Johnny Ryan**: Yes. There's another way to find that out, which I'll get into in a sec. Let me just you know what? I'm gonna just open up notes so I can write a little note so I remember. Daniel, thank you for that. As you were speaking, I have several screens in front of me, and I can see Sivan's face. And he's working on this stuff every day, so I was just checking to see if he would nod his head or not. I was gonna say whatever he said. And the reason for that is that I haven't checked the answer to that question recently. The last time I checked was for litigation in the maybe two years ago. And and the answer is that it depends. If if the company is using something called header bidding, it's a kind of a waterfall that passes between lots of exchanges. And if they're doing that on the client, then you can see certain things. But you don't get the full fast diet that that you would see if you could if you could see the the the server to server information. But what's really useful, Daniel, depending on the company, is to make a subject access request. Maybe if you're familiar with this, it it can be useful. I sent a a SAR, one of these requests, to the trade desk, for example. And I'm a person who does not drive, and I found out all about the cars that it thinks I'm interested in and car insurance. But many other things, some of which were disturbing and some of the disturbing things were accurate too. [00:42:42] **Sarah**: Thank you. Andrew is next in the queue. [00:42:45] **Andrew Campling**: Hi. Thank you, Andrew Campling. Great presentation. Really interesting, fascinating topic. Really a question for the room rather than for Johnny. What? Twelve and a bit years ago, we published pervasive monitoring as an attack post immediately post Snowden. I always felt that that was a distraction to point the finger away from us, in that case, at the government. I think this really proves that point, and maybe we need to publish a second version, a sort [00:43:15] **Simon Lerman**: of mea [00:43:15] **Andrew Campling**: culpa, and do a proper job because, I mean, this stuff is way beyond what any government could ever dream of gathering, you know, gifted amateurs compared to this. So I think it's quite shocking, and we need to reflect. So thank you. [00:43:30] **Sarah**: Yep. Thanks. Thanks, Andrew. [00:43:32] **Johnny Ryan**: I I could could I just say a word about that? [00:43:38] **Nick Doty**: Valentin? [00:43:39] **Johnny Ryan**: Very recently Sorry. Very recently, if I remember now so first, Andrew, we we we published a report called Europe's hidden security crisis. And because we liked it so much, we published another called America's hidden security crisis. And then we published one called Australia's hidden security crisis. So pick your flavor. But in that report, you'll see cases where governments are buying these data and have admitted that they're buying the data. The director of national intelligence in The US declassified a report where it's admitted that these data are very useful, And the same goes for many other states. [00:44:29] **Nick Doty**: Valentin, Mozilla. Thank you for the talk. I have two questions. So, like, how good are ad blockers at preventing traffic, if you know? And how good are trackers at aggregating the data if you use, like, different devices, different IPs, private browsing mode in your browser, and so on. [00:44:47] **Johnny Ryan**: Mhmm. Yeah. There's another question for Sivan. So I think that RTB, I think, can be blocked fairly effectively. I mean, you know, I I think you can just block it. But there may be new techniques that I'm not aware of. So people people in the room will know that a lot better than I will. On this question about cross device tracking, that's the holy grail of the advertising industry. What I didn't get into was the industry perspective, and perhaps I ought to have. So maybe I'll speak for a few minutes about this, if that's okay. Yeah? [00:45:28] **Sarah**: Yeah. Sure. We have time. [00:45:29] **Johnny Ryan**: Yep. Oh, okay. So In the in the advertising world, there's a famous quote from a marketer called John Wanamaker. Probably everyone in the room has heard it. Wanamaker, over a century ago, he said, I know half of my advertising budget works. I just don't know which half. Now that was when advertising worked. Online advertising is probably very broken. When I say probably, no one knows. We have all this tracking, but we'd actually don't have any certainty. So in the February, in particular, advertising people who are not scientists or engineers, that's not their background. They tend to be have studied humanities, that kind of thing. They started to get physics envy, and they were sold a dream that they could finally talk to the CMO. Right? So the c sorry. The CMO, chief marketing officer, could finally talk directly to the CFO in any company and give them Excel spreadsheets. So advertising, which was a kind of a a kind of a magical thing. It just paid for stuff. No one really knew if it worked, but clearly it did work. It became something that promised quasi science. And the ultimate objective of this quasi science is that everything is measurable and that you can track I'm going to make up a person, Alice. And I've seen these presentations when I was in industry. You can track Alice when she wakes up in the morning, checks her device beside her bed, and then she's listening to the radio, and your advertising campaign has has touch points with Alice all the way through the day. And then a week later when she buys your product, you can attribute that purchase to all of those touch points. That's the dream. So the dream is a omni, stazzy, we are analysis head. We follow around every day, and there is no limit to just how extensive that that is. But your question wasn't what's the dream. Your question was, are they able to do this? The answer is most of it's nonsense and doesn't work. The amount of I'm I'm gonna share my my screen again, actually. Okay. So for people who haven't yet been killed by PowerPoint, these are I'm gonna euphemistically call these bonus slides, but it's this is a second death by by power by PowerPoint. Here we go. Okay. So I'm just gonna skip through hang on. I'm gonna skip through some of this stuff about how this system is killing publishing and get into, yeah, and get into the fraud. So let's imagine Alice visits the Daily She loads the web page, and information is sent about Alice to tens or hundreds of companies. All of those intermediary companies now have a profile of Alice. They understand her as a Daily Bugle reader. That's good news in the short term for Daily Bugle. It's really bad in the long term, but in the short term, it's great news. Everyone is happy at this point. Right? Everyone's happy, Except that Alice isn't actually a person. She's a bot. Alice is is running on a smartphone in a, like, a headless browser on a on a rack in a warehouse somewhere, and she has been sent by fraudsters to go and build up a pattern of behavior that makes her look like a an auto intender, someone who's thinking about buying an expensive car. And that you know, car ads are worth a lot. So that car ad gets gets shown to her on the Daily Bugle. Everyone makes money. And then Alice is brought back. She's brought back to fakewebsite.com, and the same ad will be shown but at an enormous discount because fake website has no particular advertising value. And even for that discount, the ad tech companies are are saying to Audi or Hyundai or whoever, they're saying, look. We got you a great discount on this person, Alice, who's definitely a human being, but no one knows. And when I say no one knows, here is my plot of the industry's varied estimates of the value of ad fraud every year in billions of euro, which is very close to billions of dollars. Any dot that's gray is just The US market. So I don't want you to look at any particular figure. I just want you to look at the spread. No one has a clue what the scale of fraud is. So the dream is to to follow Alice throughout her whole day and match between devices, But the evidence suggests that that is not yet realized. [00:50:52] **Sarah**: Okay. Thank you very much for that. Next in the queue, Will Earp. [00:50:57] **Johnny Ryan**: Hi, Will Earp, SWFL. My question is that if I toggle all those boxes off, so I don't want them to track me, including the legitimate interest ones, to what extent can they still track me? That's a very good question. So back when I worked for Brave, I used to say, switch on Brave or your favorite ad browser and just click whatever the hell you want. Click accept everything. I still say the same thing. Just block and say yes or use a browser, which Braveman does, and just just block even the request. The whole thing is nonsense. There is no technical measure mechanism to prevent any party from sharing data with any other party, and that is how they make money. So in our litigation against that system, the so called transparency and consent mechanism, this is one of the points we're making. This whole charade does nothing to improve security. It's just a pain. Now it may be that when you go on Le Monde, you know, it's a reputable French website, it may be that they're not messing around. But they have several 100 partners, and there's no way they can audit them. And, by the way, there's no way the regulator's auditing them. So it's Scout's honor. [00:52:23] **Daniel**: Thank you. [00:52:24] **Sarah**: Thank you. Next is Nick, who I believe is remote. [00:52:29] **Nick Doty**: Yeah. Nick Doty, Center for Democracy and Technology. Thanks for the talk. Thanks for your very persistent work on on this topic. And I like that you're getting at at at the solution, and and I think there's a lot more work for all of us to do on that. It it seems like from from what you've shown, and and I think we we've seen a lot of places, there might be a lot of areas for improvement, but but one in particular is the longitudinal identifiers. That that that particular harm is not any individual request, any individual individual real time bidding signal. It's the ability for a a data breaker or someone pretending to be a a advertiser to collect those signals from every every every transaction you do over the course of a day or a week or a month or a year. And and I think that would give us a little more focused work on trying to identify when we introduce any of those identifiers. And so the smartphone operating systems have their ID for advertising, their or their mobile advertising IDs. Chrome has third party cookies. And so I I think those are some initial targets that should be particularly important. But I think the work that I think will be very important for ITF and w three c and other places going forward is, are there going to be other identifiers to to the extent that we can try to limit access to those unique identifiers? Are there gonna be other persistent identifiers, whether that's a IP address or maybe an email address or a phone number or other places where we might introduce user identification? Are we gonna create more identifiers that can be used in that longitudinal way? Because it seems unlikely that we're going to manage either through through industry policy or regulation, that we're gonna successfully audit every single company involved so that identifiers are not abused. It seems like identifiers definitely will be abused, and the question will be, can we restrict access to those identifiers? [00:54:34] **Johnny Ryan**: Nick, thank you for that. So as you're speaking, we had to decide when we were asking the court in this case against Microsoft, which takes aim at the these problems. We had to decide what the what should be permitted. What what is it we're asking the court to stop? And you're right. The longitudinal IDs definitely had to go, but also the probabilistic ones too. You know, I I I don't think you can have a system this is outside of IETF and and outside of your comment. But but just for the industry, like, for the for the IAB system, I don't think you can have a system that's throwing around billions of of broadcasts a day, literally billions, that are rich with probabilistic data as well. So you're absolutely right. You should you should be very careful with these longitudinal IDs. But the the for the industry, we also can't let them off the hook with things that they can claim aren't IDs, but actually kind of are. And I I just wanna give you very briefly another example. [00:55:48] **Sarah**: Sorry, Johnny. We're running a little low on time. We have one more question in the queue. [00:55:53] **Johnny Ryan**: So Go ahead. Sorry, Sarah. [00:55:54] **Sarah**: Go ahead with that. [00:55:55] **Simon Lerman**: Hold on. [00:55:57] **Tara**: Sorry. I'll I'll try to [00:55:58] **Johnny Ryan**: reach you. [00:55:59] **Tara**: Sovereign Tech Agency. I just wanted to push back lightly against the notion that 7258 is a distraction. I think if we look at today's agenda and the presentation, Johnny, you can see that the surveyor is adapting. It's the surveyor is moving off the path and onto these, like, different choke points, and I think this is a a different type of attack than the one being described by 7258. And I would welcome, like, the idea to do more work to describe the account of attack using 7258's example, not as yeah. I mean, I think I'll be careful about seeing this, but indicators showing that 7258 actually working because the attacker is moving off the path. [00:56:45] **Sarah**: Good point. Thank you for that. That's great. Yep. I'm sorry that's all we have time for. Thank you for the presentation and for your patience with the question. Very much appreciated. Thank you. [00:56:55] **Johnny Ryan**: Thank you. [00:56:59] **Sivan Sahib**: Thanks, Tony. I just wanted to say that thank you for the presentation and all the work that you do as well in this space. [00:57:04] **Johnny Ryan**: Thank you. Think it's Much appreciated. [00:57:06] **Sivan Sahib**: Really helping everyone. You might also wanna check out the chat, Johnny. There's a lot of discussion happening there, and I left some comments there as well. [00:57:16] **Johnny Ryan**: Okay. I haven't used the system before. Okay. [00:57:21] **Sarah**: So, Simon, you should now have control of the slides. [00:57:26] **Simon Lerman**: Yeah. Hey. I'm Simon, and my talk is going to be a bit related to Johnny's talk. So Johnny talked a lot about, like, how ad companies are kind of taking your data. And what I'm gonna talk about is kind of how AI, large language models kind of makes the situation worse because you cannot give, like, a like, a very dedicated investigator to everybody. And we shown a bunch of scenarios how, like, batch language models can be used for, like, large scale genome optimization. This is based on our paper, which was just accepted at Usenix. So it will be in Baltimore in a month or about. Broadly speaking, there's kind of two general scenarios we look at. One of them is a very agentic scenario where we have a bunch of data on somebody, and then we just have an agent based on a large language model. The data and uses search tools, it uses browsing functionality to essentially find the person based on that anonymous data. And the second thing on on the side, you see this the 1 to $4 statistic to try to deanonymize people. The second thing we look at is if we have somebody's identity and we use a LinkedIn account in that example, on, like, in a small social network and we use kind of parts of Reddit and we use Hacker News, for an identity, can we find the anonymous account? So you can think about it one direction as we have a bunch of anonymous data. Can we find the identity? And the other direction is, yes, so much identity, we wanna find the data from, like, a haystack of data. And these are kind of the two general directions we look at. A little bit on me. So I've done this paper on large scale denormalization. I have done a paper [00:59:35] **Johnny Ryan**: on [00:59:35] **Simon Lerman**: AI powered spear phishing. The idea there is that we can, again, use language models to search people online to find information on them. And then language models are good at writing. And given the personal information, we can then send out personalized phishing. So based on information you have online, we test it on human subjects if they're more susceptible if we give a language model access to its personalized data. And the language model autonomously finds the data. So it's an end to end autonomous process, and we find that language model get much better. It's been, like, four or five times uplift by giving them this data access. And we just put out a voice phishing study on archives. You can find them at Substack and on Twitter. We can find them at my other research as well. And we are also looking into the topic of kind of AI driven kind of that you can create on people. And I'm founding a kind of nonprofit on this kind of research direction. Okay. So going back to the paper here, this is a screenshot from our paper. Again, kind of just showing the idea how it's how we can identify people from kind of anonymous data. In this example here, this is bay just based on a real example. We have a transcript of somebody in which they were interviewed about their job position. Somebody talked about the kind of job for ten minutes about, and they already tried to anonymize it. So everything that's too are not too identifying but so when you removed, but we find there are so many details. Like, this example, an a language model can extract, for example, the field of study, they can extract this person as probably a PhD student, kind of the approximate location, and in practice can extract many more things. But you can see in some a single kind of paragraph, it kinda kinda narrow down from, like, 8,000,000,000 people. It can already narrow it down to probably a few 100 people, or, like, biology, PhD students in The United Kingdom. And so we have this extraction process, and then an agent can move on to the kind of search process. So it can search candidates online. And then it has found a few candidates that can kind of reason about them. It can compare things about their CV, some things might not fit, some things do fit. Then oftentimes, we do find it can identify the correct person. And we like to think about it in these kind of three things that language models add. They can extract this information from all data sources you can imagine. They can search and they can kind of reason about the candidates afterwards. So that's how we'd like to think about it. I'm gonna talk a bit about the kind of prior work and just to restructure it and kind of before language models and after language models. And we're gonna talk about in our precise method and results and what we can have found. Then the threat landscape, like, how could we misuse that? We probably already have a bunch of ideas how people could misuse this. Then we're talk about kind of the implications, what we could do about it, and several work we're doing on this topic. So the anonymization already existed before language models. There are a few famous examples. Two of them are from and Shmatikov. Very broadly speaking, they use kind of structured data. So Narayanan and Shmatikov, essentially, they had two datasets of people giving movie reviews. So they had the movie title, they had the movie score, the date of the review. And they found that for many people, their kind of review statistics are characteristic enough that you can identify them. So that means that, for example, they've reviewed a rare movie, but they had a rare opinion on a movie. They might have liked a movie that most people didn't like. And if you have a few of these examples, you can often identify people with a high statistical certainty. They made another paper on kind of graph structure, and they again, they found that on the kinda graph structure, so that means kind of who do you follow, who's following you. You can, again, kinda match people across datasets and across social media networks. So another way you can use structured information to deanonymize people. Then in the there were example, they were able this is essentially a forum where people got random usernames, but they were able to reverse engineer a way you got your username assigned to you, and it was basically IP address. So your username was procedurally generated from your IP address. They could reverse engineer it so they could figure out where approximately these people live. Now language models add a lot of things on top of that, a lot of capabilities. One thing is they can infer attributes about people. So before, it would have been difficult if I have a bunch of comments or post by somebody to convert it into maybe a structured dataset. I wish much what Johnny showed where he showed, like, your, you know, your gender, your political affiliation, all of these kind of it would have been hard to create a table like that just from comments. With language models, you can with, you know, with some quality, infer all of these attributes from free text and all sorts of data you can find on people. The I think the language model's add is they can just and it performs all sent on people. It's open source intelligence. So they can go online, check websites on somebody, and find information on people. And there's a paper I was involved in, and then we use it for spear phishing. I mentioned it earlier. And then other people have also tried out agentic identification. So giving an agent some data on somebody kinda to find them. This will lead, try to have small scale of people. And we essentially try to do this in a large candidate pool. In some cases, we have ground truth. I can describe how we got that. And, yeah, we showed that it's it's very capable in many situations. Okay. I'm just taking a step back here. Why is it important at all? So there are many people who have some kind of account online, probably most people in the audience, perhaps on Reddit another social media platform. Maybe they share some of their opinions on there, maybe on politics, maybe on the hobby. And oftentimes, usually, people don't use their real name there. And oftentimes, people can assume they are among you, miss. And they might say some things that might just get them to trouble, or they might say things that they just don't really want to be tied to their, you know, identity publicly. But I kinda assume that by just kinda writing a pseudonym a pseudonym, you know, fake identity, they cannot protect it. And we kinda showed this is much more shaky now. And one way to think about this, in the past, there have been people like Satoshi or Ross Ulbricht who were, like, really high value individuals. And there were, like, thousands of us trying to find them. And in the case of Satoshi, they kinda narrowed it down to, like, a few people. They're still not quite sure about who he is, but Ross Ulbricht was eventually caught. And the idea kind of is if there were, like, a 100 people trying, like, crazy to find your anonymous accounts, maybe you can intuitively get it would be kind of dangerous. It might eventually catch you. This is kind of what's happening with AI. You have they get very cheap, and it's very easy for people to kind of essentially investigate people with language models. So these are our results regarding the agent identification. But this is we give some data to an agent that expects information that searches the web for this information, tries to reason who's the correct person. So we had the the topic interview dataset. This is a real dataset that was really anonymized, and these are these job interviews that are 10 long that so you get Anthropic or some involved in releasing them. And there are nine out of a 125 who were able to find them from a ten minute interview. That was anonymized already. So they already, like, redacted parts of it. And then in in our test, which is a it's a bit synthetic. So out of 338 so for these people, we knew their identity, but we essentially redacted the the information. And then we gave that to an LM agent, and we saw, can we identify them? So in this somewhat more synthetic setup, we're able to get 67%. But it usually depends how much data you give them. Ten minutes is not a lot of text that gets generated. But you can definitely see that in many scenarios, you can actually identify people just from kind of anonymized data. Yeah. And and just it counts depending how you do it between one to four units deep. And this is kinda brought you what it could look like to identify somebody. So in this case, the left, this is some kind of forum, and this is a cybersecurity specialist. He kinda talks about his job. He talks a bit about where he lives. And in this case, Swiss cybersecurity researcher. And at some point, he mentions here that he gave a talk at the security conference in Switzerland. And our LLM agent sits sits on the right side. It takes this detail. He gave a talk at this this security conference because that narrows it down quite a lot. Like, searches for different conferences in Switzerland, searches for talks that match the interest of this person. And from that, it finds three speakers. It searches over them. It looks who is kinda the best match given all the other information. At the end, it finds a LinkedIn profile of the person. This is what it kinda looks like. So the LLMs find some kind of smoking gun evidence so that, in this case, you have a target of security constant in Switzerland. So that nails it onto a smaller set of people. It searches over these candidates, and then it can reason over these people who matches the person the best and then finds that individual. And then well, again, that that's the agentic identification where we have permission on someone and the that case is those personal forum. We wanna find a name and identity. Now we do the opposite. We we assume we have a real identity. We wanna want to find the account. So these are synthetic again. In this case, we actually use the language model to run over this form and remove identifying information. And in this case, this is how we essentially got to ground truth on them because we knew who they are. Then later on, we anonymize the data in the forum. We quite ground truth, but it's a synthetic evaluation here. But there are some distribution that biases here, but this is where we get much more data and ground truth. So, again, we cannot think about this in this kind of extract, search, and reason framework and calibration depending on how precise you want to be versus how sensitive you can calibrate given different methods, which we described in our paper as well. That's how we kind of like to think about this. Yeah. So this is about the calibration part. So, again, we we we extract data, what people we search over candidates and choose the best candidate. One thing that's kinda missing there is this kinda calibration. Sometimes you want to be very precise. You want to be very certain you got the correct person. Other times, it doesn't matter so much. And we test kind of two methods of calibration. One of them is that we simply ask a high reasoning language model to tell us how confident it is. These are the same person as a correct match. And then given that confidence, we can essentially bury a threshold to set to to trade off precision versus sensitivity. The other method we use is, like, an alternate sorting where we show an LM two matches at a time, then we ask you which is a better match, which is a bit more complicated. We actually found that language model are pretty decent at just giving us the confidence. So we must just use the first method. We just ask a language model, how confident are you at this decision? And so in the past, again, they just use the structured data. But now with language models, we can also just look into the visual review text itself. And we can, for example, look at the fact that two people describe a scene in in the same way, and that's strong evidence. So the open evidence, I could only look at the numerical score, the movie title, then we can look into the movie review, and we can say, you know, both these people enjoyed this one scene, so that makes it more likely they're the same person. And this is so these are the results on one of our tests. In this test, we had people's LinkedIn identity, so we knew the identity. We were trying to find the Hacker account. News is kind of a popular online forum, which shows some of you have used before. And did you ask for how many leading users can we find their Hacker News account? And we actually find that we can if if we have a precision of 90%, we can actually identify 55, 54% of people corrected. For higher precision, 99%, we can see it's 45%. So we can get a very high precision asset sensitivity at the same time. And we compare that to these older structured methods. We found that these work much best. They only get about 100,000 people, whereas we essentially get half of people here. And the plot you see on the right side so one question people have is, well, Hecan is a pretty small system network. It only has 89,000 people on it. We find if we reduce the amount of people, we start at a 100, then we go to 10,000, 90,000. If we extrapolate it, it should still work decently well at a 100,000,000 people. We should still get about 30%, a bit less than that. This is about the Netflix movie tech. So we wanted to make a direct comparison to movies. The way to read the plot on the right is essentially the area under the curve. If more area under the curve, it's better. And the blue line is using our large language method, and the red line is now in none. The old structure data method, again, you can see that it works much better even on the main of movie reviews. And this is the another kind of scenario we tested, And this was a we call it a temp power split. So, essentially, we asymptotically took a Reddit user's activity over three years. We took the first and the last year, and we pretend it's the two different accounts, and we saw if we could match them again based on the activity. And, in red, you see a method similar to now in none. But in this case, the yellow line is using the strongest term of our method. So, again, there's a much better area on the curve using our method. Okay. So I'm gonna talk a bit about how people could misuse that and what are the implications here. Think language much in general allow much more mass surveillance, and they can, like, go through enormous data sets and things that are supposed to be anonymous given the effort of language mods. I think they will often be able to identify people. There are much more opportunities with language models. Just in this way, they just kind of come to all of the data and to make instances that you would need thousands of people for in the past. And, of course, you could also use it to detect crime and terrorists in some cases. And this similar to what Johnny talked about, could be used for more advertising and cross linking of devices for advertising and duplicate through scans for more social engineering, automated social engineering. And you could try to, yeah, identify key employees and decision makers. What can people do about it? Big platforms should be careful on which kind of data they release. If language mods can come through it. Again, just really protect you from kind of, you know, third party actors who may wanna misuse the data. AI labs can try to monitor for this type of misuse, but there's open source models, so I just have I think if there will be people can be more careful on that. Yeah. And I'm I I teased it up in at the beginning. So another thing that people can do with language models increasing the kind of AI just years. So what this means is AI can kind of come through all of the data. It can try to find here anonymous account. It can try to just take out a public stuff that's there on social media. Maybe your friends have posted about you. The stuff that's maybe on a company website. Also, breach data from hacked websites in the past. They can put this all together in a way that the possible of taking a human investigator that can create this AI dossiers. And you can see a tiny, accept on an AI dossier on myself, and it found an email and a outreach dataset. It can found, like, my Instagram account. It's found where I live and many piece of information on me. And language mods enable a lot of new kind of misuse with that. Yeah. And I've talked about, like, all sorts of kind of misuse scenarios about, like, spear phishing or, like, voice spear phishing, and they can perform more and more complex tasks. We have this meter graph, which most people have probably seen. The systems can perform, like, sixteen hour tasks, and it's rapidly increasing, stabling every four months, let's say. So the potential for miss is is getting much, much greater and the potential for kind of things going catastrophically wrong. And yeah. I mean, there are serious scientists who think that eventually AI might resist our control altogether, and it will be harder and harder to keep the systems under control, and the kind of failures will become more and more impactful. Yeah. As a kind of conclusion, we find that the classical methods much worse than using language models, and there's much more new things that language models enable. We think it's where our method scales to larger datasets. And for just a few dollars, we can identify individuals. So I think a lot of threat models have to be reconsidered. A lot of kind of data privacy has to be reconsidered given kind of the new threat that touch language models post. Yeah. If if you're interested in this kind of AI, can also send me an email about that or go follow me on, like, Substack or Twitter on this topic. And you can scan the QR code to see the paper on the topic. Yeah. [01:23:10] **Sarah**: That's great. Thanks very much. We have time for a few questions, if there's anybody who wants to join the queue. Tara? Yep, please. [01:23:23] **Tara**: Tara, Sovereign Tech Agency. So if I'm understanding this correctly, it's it's not discovering a new type of linking people. It's just using the already existing exposure but doing it a lot quicker and much more efficiently. Have you looked into how it compares against just random linking, let's say, like, a casual account versus someone who's actually trying to hide? [01:23:55] **Simon Lerman**: Yeah. So it's very difficult to make a dataset of people who are actually trying to hide the identity online. I think that the Anthropic interview dataset which I discussed earlier, that's pretty close to that. I mean, they try to redact the interview transcripts. So these are not necessarily people who went out to prevent being identified, but the people who made the dataset tried to hide people's identity. And we are still able to get some of them despite these measures. I'm sure it would have been much easier if they hadn't done that. And it would have, like, mentioned the university name that doesn't type of information. But, yeah, I I think even if you try to hide some of the information, we can still identify you in some cases. [01:24:48] **Sarah**: Andrew? [01:24:50] **Andrew Campling**: Yeah. Hi, Andrew Camping. Just reflecting on the two presentations helpfully curated back to back. Thank you, chairs. You can only imagine the number of data points against a profile from Johnny's presentation feeding into the language model. You it looks like it'd be trivial to de anonymize many of those. [01:25:10] **Sarah**: Exactly. The concern that's, yeah, running through my head as well. I think Shivan had a question. Last quick one. [01:25:21] **Sivan Sahib**: Yeah. Thanks, Simon, for the presentation. Just wondering, did you also use stylometric identifiers, or was it only semantic in your analysis? I think when I read the paper, I thought you had said that you don't use it. That's, again, for further exploration, but I thought in today's talk, you said that you did examine, like, your writing style and things like, you know, certain phrases that I that an other that that an author might use. [01:25:52] **Simon Lerman**: Yeah. That's a good question. And so by and large, it doesn't use telemetry. I think that the language model does have the ability to look, just for example, for example, at the different movie reviews, or it looks at the interview transcript, and then it searches the web and might find some other things you've said online. We never told it, oh, look at the stylometry. But the language model can certainly without us telling us, it could still look at this type of evidence. So if if you and that interview transcript use a phrase and then a fancy phrase, I just assume a standard type of speaking on your other personal website, but maybe finds a different transcript for a different interview. It could use that information, but we never encourage you to do it. And we we prefer, like, semantic data because it's also easier to verify for us for the language model to be certain, you know. If you say, I am a student at university x doing y, and we find on LinkedIn just just some exact information. We can, like, verify it ourselves. That's what it's mostly using. Yeah. K. Thanks. [01:27:14] **Sarah**: Thank you very much, Simon, for that excellent presentation. I'm not sure many of us will sleep quite as well tonight as we had if we didn't know most of this. So, thank you for your input. [01:27:28] **Sivan Sahib**: And I'd like to add that one reason we did choose these two together is to see if people have some creative thoughts on on what communities could do to improve our situation. So, just it occurred to me when you were talking about the people who are trying to who hide and are harder to find, whether there might be something like like the things they do at DEFCON, but it's a hide and seek game. And, you know, incentivize people to to really learn what the what the parameters are and what their their risks are, but creating new techniques for hiding as well. But this just we'll ask this on the mailing list, but we'd love to see some some blue sky thinking about community responses to the kinds of things that that these these two wonderful presentations have have shared with us. So thank you so much, Simon, and retrospectively, Johnny as well. Really appreciate this a lot. [01:28:33] **Sarah**: So we have a couple of minutes left, and I'm going to apologize to Gianpaolo for squeezing him in. So, let me give you the slide quicker. Okay. And there we go. So, I'm wondering if you can, in possibly two minutes, give us a lightning summary as a precursor to a discussion that we'll take to the list. [01:29:01] **Gianpaolo**: So this is an analysis of what is happening after the ECH deployment. And the fact that ECH is called the end to end encryption, but in reality, it is hiding the destination while the source privacy is not touched. What has happened is that from HTTP to HTTPS to HTTPS to plus ECH. Now we have a destination hidden from the casual intermediate people. But now we have an high priority element that can do linkability of source. So we have the client facing server that has the ability to do profiling over billion of users because after the deployment of ECH, what is happening this is my analysis over 1,000,000 domain, is that one single ASN is managing 33% of traffic. Top one IP, 4% of traffic, and top 10 IP, 36% of traffic. This means that while in the generic observer now are weaker because they have more visibility over the destination, these privileged observer, the client facing server, has this very strong linkability capability. So has the capability to create profiling over billion of users because it's the front door of the CDN. So ECH is not bad. It's good because it's providing half of what is needed to hide the traffic in Internet. But what is missing is a counterpart to protect the source. Because if we have only the client facing server, then we are creating an a privileged observability point that can do profiling over the whole humanity. Just one actor has access to 30% of Internet traffic. So the privacy is not disappeared, but now it has been displaced in a privileged point. Also, in if we are using CGNAT, this is not decreasing the capability because CGNAT is quite static. So by attributing one IP address for a longer amount of time, it is possible to create profiling. What can happen? I'm doing this pineapple pizza scenario by creating a crazy owner of a CDN that high hate pineapple pizza. It can create a profile of lover of pineapple pizza, increasing prices, for example, slowing down services, and the pre create exclusion of additional services. If this CDN can also provide identity services, then the anonymous profiles can be enriched with the identity. And we are creating, as we have seen with the previous presentation, a profile that has a very high visibility. So what I'm proposing is to study what can be a method to have the companion of the client facing server on the network edge to anonymize and randomize the access to Internet to decrease at least the capability of client facing server to do profiling. So to reduce the continuity of the source IP address. This is a hypothesis, but also I think there can be other methods such as doing the discovery of client facing server over different endpoints or for making it not easy from browser's perspective, the usage of the same CDN for accessing a huge amount of domains. So the ask is if this is so source privacy is an important point and in case to continue study on how to improve the source privacy. [01:33:20] **Sarah**: Thank you, and thank you for the speed. Appreciate it. So I think as chairs, we will take an action to open a discussion on the list to frame what specific acts aspects of this will fit into our charter, and what pieces of research we can bring into the group on this basis. I think thank you for bringing the topic to the group. Appreciate that. That concludes the PERG session for today. Thank you everyone for participating, and we'll see you at the next meeting. [01:33:58] **Sivan Sahib**: See you. Bye. Sorry?