Decentralized AI involves distributing data processing and model training across multiple devices or locations rather than relying on centralized cloud servers, offering benefits in privacy protection, data sovereignty, and reduced dependency on single providers. Key approaches include federated learning (where models are trained locally and only updates are shared), on-device inference (running AI models directly on user devices), and distributed data governance (decentralizing data management responsibilities). This approach addresses critical challenges in modern AI systems, including privacy concerns, data security, regulatory compliance, and the risk of monopolistic control by large technology companies. Organizations can implement decentralized AI through local-first design, using open-source models, and establishing trust-based data relationships, though challenges remain in data quality consistency, coordination overhead, and infrastructure complexity.
Decentralized AI: Data, Governance, and Edge Personalization
Added:Hello there everyone.
Hugo Bound Anderson here. So excited to be here today with Katherine Jamal and Joe Ree. What is up, Katherine? And what is up, Joe?
>> Not much, man. I just got back from a hike, so if I'm looking a bit sweaty right now, I uh I was up, now I'm back down, I guess.
>> You look fantastic.
>> Thank you.
>> Um and you're in Utah, Salt Lake City.
And Katherine, >> you are in Berlin.
>> This is correct.
>> Fantastic. I just >> not not much going on here.
>> Well, I I I don't believe that for a second about Berlin. Um but we're here today to talk about decentralized AI.
Um, and for everyone watching in the chat, I've put um we'll be having a Q&A discussion on on Discord. If you want to join the Discord, um, I can then filter the questions through um, and it'll be in the podcasts channel. Um, so just super excited to have you all engaged as much as as much as possible. Um, I this is a live stream of the Vanishing Gradians podcast, which we used to call a data podcast. Then we called it a machine learning podcast. Now we call it an AI podcast, but of course it's data all the way down as as we know. And we're here to talk today about the fact that the AI revolution will not be centralized. Hat tip to Inis Montany and her talk, the AI revolution will not be monopolized. Um, as Katherine says, today's breakthroughs in hardware, privacyware, machine learning, distributed data, and compute, uh, show us there's no reason to upload all of your data to a single model provider.
These are the types of things we'll be getting into today. Um I just wanted to say give a brief introduction to um uh Katherine and Joe people who need no introduction um but Katherine has worked in the space for a long time and offers trainings and advisory focused on building privacy first data and AI and ML systems um products and platforms um definitely check her out on on LinkedIn and her website is it still um kjamistan.com Katherine great um she's also author of the O'Reilly book the wonderful O'Reilly practical data privacy. Um along with an earlier O'Reilly book on uh natural language processing in Python, which is what originally brought us together. Um Joe um does many things. He helps data teams build systems at scale with clarity, real world insight, and a healthy dose of skepticism toward hype. Um and you know, the healthier the dose, the better at the moment. Don't forget to eat your data broccoli people with with Joe Ree.
Um, Joe is also never said or thought that before. Joe is also co-author of the bestselling fundamentals of data engineering um on O'Reilly used by hundreds of thousands of professionals worldwide. Um, is is there anything I missed or got obscenely wrong in that introduction, my friends?
>> Crypto influencer. I'm kidding. Um, >> well, someone called me like a content creator the other day, and I I I do I think my my my therapist is definitely going to, you know, get a few vacations out of that characterization. It's something I'm I'm I'm reckoning reckoning with. But look, why don't we just jump in? Actually, no, let's I I just want to say by way of motivation, the reason um this podcast live stream came about is Joe and I did a podcast uh earlier this year where we just started talking about our mutual friend um Katherine Jamal, who we both learn from so much all the time and love dearly.
And we were like, we should do one with her on decentralized AI. So, we pinged her and actually I spoke with Katherine about it recently in Berlin. We started a group chat and um here we are. So I'd love Katherine if you could just help us set up the conversation. So to have c decentralized AI, I suppose it's worth talking briefly about centralized AI. So what is centralized AI and data and what benefits does it have and what problems arise from it?
I mean to some degree there's no such thing as centralized AI anymore in the traditional perspective because obviously as Joe book shows as Hugo your work shows um we don't even keep all of the data and the processing in the same place anymore and haven't for quite some time. So to some degree it's like breaking down the myth that we do have like this central monolith of some sort of software central monolith of some sort of data and central kind of central guidance for all of that. And I think that um once we start to rip away the fact that we already have uh quite a a distributed or decentralized compute environment and reason about it that way, then we can start to actually decide um how would we like the system to work in regards to perhaps other concerns we might have like privacy or security. And instead of just buying whatever is easiest off the shelf, actually start to reason and think about the way that our data moves, where our data is located, and then therefore also where things like our training or inference runs and how we can set up the system to actually reflect um not only cost cutting mechanisms that we might want to optimize, but also things like obviously my work focuses on things like data privacy.
And so what what type of benefits do I suppose the what we consider to be centralized AI have have now and then why are we here today? Why is thinking through decentralized AI and data so so fundamental for us now?
>> Joe, you want to talk about uh the myth of the growth of the data warehouse?
>> How how much do you want to >> Well, I I don't think that the data warehouse goes away. Hey, I mean the central notion is, you know, integrating data from across sources and, you know, having that in a common place to query like that's, you know, been been here for ages. But it's interesting now because you mentioned a couple interest um things that I'm curious about and I've been thinking about myself, which is, you know, you mentioned things like figuring out your data and getting it in order and all this other stuff. And I'm like, okay, so why weren't we doing that before? like what you know is this my my biggest question right now is this time going to be any different you know cuz this is all the foundational hard work that we're putting into it this is something I uh podcasted about on Friday actually where I talked about um Satia Nadella's letter to employees whoever happens to be left at Microsoft as well as um you know uh shareholders and just you know pundits and so forth but you know he mentioned a very interesting thing which I've been stewing on been doing thinking about a lot which is you know there's been a lot of capex from a lot of big tech into AI data centers you know model training etc etc etc it's unclear if that money is being recouped yet right so like it's very unclear we're still running at a massive uh loss I mean the um I think the definition that openai and Microsoft had was for AGI was hitting 100 million in gross profit or something like that uh from AI and at that point you would say AGI. Okay. So so the my my biggest question is you know we're putting a lot of money into this stuff. Um the uh success rates have been quite abysmal still enterprisewide like production implementations of AI. So and if you trace back a lot of it's data a lot of it's the same stuff we've been dealing with for decades. And so I'm kind of like, you know, is this the moment when we're actually going to start working on better data, working on better security, working on better this or that or the other thing, or are we just going to punt it again like we have with every other um you know, craze that's come along? Uh which there's been many. So, you know, it's it's one of my big questions in terms of getting to any type of production AI or the central questions. um or as a supposition that AI will just become so sentient and powerful it'll just solve all those problems for us and then we'll just be on our way right that is a hypothesis that some people share so you know I think that what your your central point which you led up with I think is a very interesting one because it's it's one that we've been struggling with as an industry for decades um not just today but um you know and so it's it's an interesting one like we we're still barely getting data warehousing correct right now we're going to other stuff, ambient agents, all this other thing. All right, so this is this is the big disconnect I've been thinking about lately. Uh, but the messaging for big tech is you got to do this because we spent a lot of money on the and we have to like get paid somehow. So, you know what I mean?
>> Yeah. Yeah. I mean, I would argue that I mean, we I don't want to go on a tangent. I would argue that the recuperation comes when targeted ads come with your with your uh monthly pro subscription, but that's neither here nor there. I think I think your point >> actually a pretty big point, but yeah, we can talk about that later. Um, I think your point though like going back to the idea of like centralized versus decentralized is kind of related to like a continued argument I feel like in data governance circles and I would love to hear your opinion on this is um a big push that I have as more like a privacy security oriented person is perhaps like we would need to have less governance a if we collected less data and b if we allowed what is like the true meaning of data protection laws, which we're not going to get into the actual what got implemented or whatever, but the true meaning is like there's a relationship of trust between the person giving the data and the person managing the data.
And therefore, things like data quality are easier to guarantee because there's like a a direct relationship there. And then secondly, perhaps not all of that data needs to be you know um centralized in whatever it is warehouse otherwise and if we're actually going to do this I agree with you by like principles by the original ideas let's go back to you know original design of computing centers we might have governance experts then that are either domain focused or product focused or something like this and they can help guide some of the way that data comes into our our centralized spaces and then also in directly interact with people who might end up being able to help manage their data and the veracity of their data if it is available to them and made visible to them. Now, that hasn't been incentivized by any means because most of the ways that we've made money from data have been by making it quite intransparent the way data is collected and used. And if that continues, then we I agree that we will continue having the same quality governance and I know your special interest modeling problems because we're just like slopping up whatever we can get. We're storing it in whatever quoteunquote idea of centralized that we have and then we're just giving access h and h and um anybody can grab anything.
Uh what I challenge us to think about or I guess like why decentralized AI is so interesting to me is perhaps we can think of leaving data closer to the edge or at least you know closer to the edge than we have and therefore have not only better trust and privacy guarantees but a better conversation with the people is who data is actually is so that I can go in and say you know often my my Google ads thinks I'm like a 35year-old man I think you know what I mean so I could So too actually I >> Well, there you go. But you know it's um like perhaps with the companies you have trust with then they can have better data and um yeah be curious to your thoughts.
I mean it's it's an interesting one right the decentralization argument's been going on you know it's it's been I mean crypto you know that was when I think in tech it sort of blew it open you had microservices before obviously and few other attempts at decentralization with data obviously with this you know this this sort of blew the lid off on the data world and now you know Jamac's working on a very similar or extending the architecture of data mesh into you know autonomous agents and autonomous data products and So you know and it's it's been an interesting journey I would say if I look at what's happened with well there's a couple threads I'll talk about data mesh first I'll talk about federated uh data governance next which is an interesting topic but the um you know data mesh was uh it blew onto the scene what kind of the late uh 2010s with two articles that um that did on thought spots blog and that ushered in a whole revolution which I I still am very bullish on. I think it's just going to play out. Um, you know, the challenge of it is just that organizationally, right, a lot of companies are still it's just classic Conway's law, right? Like you can't get around it. It's just the universal gravity of like an organization. Um, so you're going to build the systems of communications and architecture according to the org chart and um how the company communicates and so forth. And so that's then sort of one of the the hard things. And then I think also decentralization became sort of a um what what I like about what you're saying is there's there's a notion of decentralization I think in its purest sense which we've you know been I think sort of I remember the early days of of uh the internet and the web um you know growing up with that whole notion of this kind of global village and everyone you know owns their data and and basically has a got you know homestead on the internet and so forth and like that was a I think a an interesting vision back in the early 90s um that I subscribed to as as a you know as a teenager and I still wish that would have happened. Um then along came the corporations and the big money and then that um obviously changed because you had the walled gardens like you early on computers AOL blah blah blah they got you know blown apart by Google basically and the web browser um you know but that ushered in basically other types of w gardens like Facebook and Google and everything else right so but it it's it is interesting I think in the sense where if you couple that whole trend with where companies are um there's been a lot of parallels um you know I think the decentralization like everyone was interested in data mesh for a while right I think they still are but I think at the same time they started looking at this as um you know data sharing for example became or data products became basically I'm going to share tables or I'm going to share uh you know an Excel file in SharePoint or something like that that's not quite the same because what's missing is that federated computational governance layer that sort of underpins everything and that's still being built I know that's what working on and So, you know, but I I I do wish that we would get to that notion of decentralized data. I mean, I've, you know, consciously made it a point to use um, you know, privacy browsers. Now, I don't use Chrome unless I'm absolutely forced to. And, you know, doing my own little things there and, you know, it feels good to give a you know, subtle middle finger to big tech.
Um, you know, but but it is it is fascinating because I see I see the same thing happening with AI models right now, right? where like the the creation of these models is concentrated in the hands of a few players. Just it reminds me a lot of of the walled gardens that we saw in the early days of of the web.
Um and the the information superighway where you had like to me open AI is basically AOL of today.
Like it's it's almost literally the same thing like it's you know for the longest time AOL was synonymous with the internet. um you'd get those CDs in the mail and you know I'd use them for coasters for for beer and stuff but like that you know but but there's going to be something else that comes along and I think the interesting thing is especially right now with sort of the way geopolitically the world's going is becoming a lot more fragmented and I think a lot less certain in terms of which direction things go. So I think this is a good time to have a decentralization discussion and sort of think about what's next and data governance circles. I'll just finish out here. I'm working with one author right now on a book on um federated data governance and that's he's taking a much different approach. Data governance up to now has been much very much a centralized committee. I would say you know um it's almost become its own meme in some ways like you're doing seances over the the data and trying to govern it and and so forth. I think it's mostly been a lost cause um in a lot of companies. It's it's seen as basically a barrier to everything. So people just circumvent it and it's easy to do because you're just a committee that doesn't get a lot of attention normally.
Um often seen as an annoyance. You're like cool I'll just go behind you and you know do everything and I don't need that's why you have shadow data projects everywhere. Right. The interesting thing right now is with vi coding apps um you know I'm very curious to see like what this next shadow um it looks like when everyone's just voding apps galore letting them run the business and nobody even knows that they're there. Uh this will be very fascinating to see. Um so yeah, I I think there's a lot of um things, but the the decentralization argument I you know, I'm very curious to see how this shapes out. Um I think we're very uh uh like-minded uh in a lot of ways, KJ, and and probably uh um Hugo as well. And and I don't I personally don't like the idea of everything being controlled and homogenized by big tech companies. I think it's it creates a very boring society. Uh and I think it's really dangerous.
without a doubt. And um I I love that you named data science. I haven't heard that that term before, but you know these acts of of naming. This is something I'm getting tingles down my spine even should even maybe organize a group called data science because a lot of it has been incantations for for so long. Um I also maybe we can move the conversation in this direction at some point. I I do love that you um mentioned a um open AI is a AOL and I have I' I'd be interested in chatting about what we all think the moat for open AI is because I kind of half joke is that they don't even have the network effects of of MySpace. Um >> is lobbying.
>> Well, right. Um and and they also say memory is a but I don't necessarily buy that either. I um Laura Summers has a wonderful question or or series of interconnected thoughts and questions in in Discord. Laura asks, um, this is something I struggle with. Why is it so hard for people to conceptualize decentralized data? Is it just too much given the challenges of centralized data? A lot of the problems we were describing in data warehouses or do we need to see examples of success stories before any of the big corporations will be brave enough to try this out themselves?
>> I don't uh I don't think decentralized data is incentivized if you're already a large tech technology company. like by default the moat of Google let's say when we say we we're going to Google something which was for years and years the source of information right um the mode what there was information mode and also the ability that everybody used Google is even more information and I think that's the same playbook that open AI is trying to play except um the most interesting new development I've seen with open AI is if you look at the more recent papers like half the authors are writers.
They're like hiring like sitcom writers and other writers from Hollywood and stuff. So I think like the idea is that openi can create personalities that you find engaging and also be your single source of information. This combination your be your new AI best friend telling you what the truth is is perhaps uh quite valuable if that actually comes to pass. But decentralized data I think like Joe says and I don't know I we are being very millennial right now but like as also a '9s kid on the IRC chats and with web rings and this and that. I think there there's been a big push since that time against having extremely diverse voices given fair audience play.
And what we've seen with the growth of influencers and so forth is like people have been trying to do this but it's always on some particular platform and it's always with some particular angle angle and of course it may or may not you know be using misinformation or other types of um things that perhaps we actually as a society don't want to some degree because we we gain these engagement metrics. But I think that the hard part of thinking through decentralized data is um what I what I referred to in a keynote a while ago as like the vycal moment for decentralized data. In that case I was talking about personalized AI. So what how am I going to personally use AI for myself locally?
But I think what we need to see and the visi was like the moment that the Apple computer took off. It was like the first time that somebody on a PC saw that they could use their data to do something interesting. It was basically like an early Excel program. And what I would love to see and is what is like the Vycal moment for using your own data on your own device in a federated manner that you're going to benefit from that makes it then more valuable for you.
It's like the web ring space of today then contributing to open AI or Gemini or whatever it is. Right.
>> So from a user perspective then I'm just wondering what a decentralized AI experience would look like because as we've said there's no need to now upload all of your data to a single model provider. So what what's the vision then for decentralized AI?
I mean I don't necessarily think there has to be one vision. I mean so I come from a background of federated or distributed learning. So this could be the case that numerous people, numerous organizations or entire community or groups of friends like the three of us can get together and train something together. So that is perhaps one view in that we don't need to put our data all in one place. We just need to organize that we exchange you know updates and and other things as we go along. But I think that that's also like a very myopic view of what that's like one of many things that um it could look like.
It could also look like I think I run four models myself locally on different hardware that I have available and I build myself kind of super super system that I use where I send out different queries to different things to process, right? Or it could be that Joe and I Joe lets me use some of his hardware or access to some of his things so that I can build something for him, right? It could be any of these things that doesn't require one centralized essentially manager of all of these which therefore you know is why it offers an alternative to like the tech powers that be.
I was a bit disappointed in Apple really because I thought that they they would at least have shown a model from a uh a scalable hardware perspective of of you know having federated um ondevice AI but I it feels like they all but gave up on um you know that that approach. Uh I I know people at Apple and I think you know one of the one of the people who work in data at Apple and one of the big problems there is there's no data to work on. So like they've been trying to build machine learning models on um you know for various things and the problem is um lo and behold you kind of need data to train on I think they tried to do synthetic data for a bit and that didn't work too well. Um in fact it's interesting so the only data that Apple can get from you uh is if you opt in to their uh longitudinal health research app and even then it's really locked down.
So I, you know, people I know there, they spend most of their time talking to privacy lawyers at Apple and they're usually, you know, told yes or no, you can do this or can't. Usually same monotone fingerpointing I just gave. Um, so that's kind of the challenge, but I was really hopeful that Apple would have some because I know they had like the ondevice model that you could use and that that would be pretty dope, but I don't know. So it's uh anyway so I don't so be curious to see like what what approaches we take you know to to make decentralization a reality. Uh so you know cuz right now my my thing is you know for kind of just whatever stuff whatever type of work I'm doing coding I use claude Gemini I sort of just use for like regular search you know if I need to because it's Google's data and it's pretty good. opening I just use for like general questions because it does have memory and it does really gaslight me in terms of being a good friend. Um but if I have you know private questions or stuff like that that I don't want being ingested all I use just hosted models on my laptop. So, O Lama for example, right? I just asked that and that's um you know presumably private but you know Sam Alman the other day he talked about how yeah anything you type into chat GBT is discoverable legally.
So like that should give you pause like don't use this thing as a um you know for your deepest darkest uh you know secrets and stuff. Um, but yeah, so it's it's interesting one, but you know, it's interesting when I talk to my kids, too, cuz they hate AI actually. They're like 12 and almost 15 and they uh like, yeah, it just because I think what they what they see is it just makes everything the same. Like they can catch on to that, right? So that was pretty interesting.
>> Yeah. And I think Joe, like you, I was pretty bullish that Apple would do something to release some foundation models that you could then personalize.
like Apple also hired some really interesting researchers in the space of taking local data sets and then essentially personalizing via some of the ways that the embeddings work what was happening on your device like data about you or local data as it was called with some type of idea of a foundation model or something like this.
But I was also really disappointed to basically see the message uh we're giving up for now. Um, I think it leaves a big open playing space for potentially organizations that have a lot of data like let's say publishing companies or media companies or something like this to pair with consumer grade um types of companies like Apple because in my ideal world you have a foundation model that's fairly trustworthy.
Well, you you can maybe choose whatever flavor you have and then you have the ability to take your own data and easily um transfer some of the ways that you like to consume information to the models for your own personal experience.
One of the ideas that I had around that too or one of the things that I would love to see is then that I can compare my personalized model with other people around me and decide who I might like to either exchange models with or train with um because we have mutual interests or because also our interests are maybe divergent and maybe that's also useful for me to hear divergent viewpoints of something like this. I think there's opportunity now that the information game is becoming even more consolidated that that may offer an alternative, but I'm I'm not sure.
>> An interesting idea though, sort of this like meta vector map of like everybody's vectors and then we uh how similar am I to KJ or Hugo? Um yeah, it's uh and Google kind of did I mean, it's interesting is Google had the ability to do that back with Google Wave back in the day when they're trying to break into social networking. Um, they had everyone's profile and interest and they uh, you know, in typical fashion didn't do it. But, uh, yeah, it's interesting.
>> I miss Google Wave a lot. And in fact, this is this is going to get really geeky in along in a different vector now, but I was actually at grad school doing my PhD in math maths at the time.
And the the renderable latte in Google Wave is how like I shared um like manuscript stuff with collaborators. It was that that was that and like macro you could define your macros in there and stuff. And then, you know, once again, killed killed by Google, um, which we'll definitely link to killed by Google in the show notes. Um, I'm interested, we've talked about a lot of kind of different interesting aspects of decentralized AI, but for viewers and listeners who don't necessarily even know what, you know, federated learning is, for example, or distributed governance. I'm I'm wondering, you know, if you like both of you could kind of just give us a 101 or or beginner's guide to kind of the important important parts of the space.
AJ, you can take it off or kick it off.
>> Okay, we'll start think about this a lot more than I have. So, >> uh just in a different way, I think. Um but yeah, federated learning uh let's take that federated learning came originally from I mean probably was started before that but the first time that it was really published that it was in production systems was from Google who was using it essentially to do next word prediction for smart keyboard because it was noticed that for example predicting non non-English languages uh so languages other than English next word prediction and particularly things like emoji use and other things like this was not very useful. Obviously, collecting keyboard data is super icky and you should make sure that no apps on your phone collect keyboard data. That's like pretty bad cuz then it's your keyboard for everything that you type into. Anyways, that aside, um they were like, "Yeah, we probably shouldn't do that. We're probably going to get a huge fine if we do." and therefore developed the idea of federated learning which is essentially um let's push let's decentralize the learning process. Let's push the learning to the edge. So learning can happen on device and then only the updates which either means the gradient updates or the weights of the local models or some abstraction of the weights of the locals models are then exchanged um usually via some sort of uh centralized or regional aggregators that are then used to update the models with an average of all of those updates or with some dropout variant of those updates and then sent out to continue the processes. Now that was um I don't know about 8 years ago now. So obviously the field has gone much further. There's a lot of really interesting things.
Flower is a group um that was originally a research group and now a product that recently trained across several continents. The first LLM trained in a federated fashion. Um which is cool because I feel like they're breaking down barriers. A lot of people are like but latency but this but that. And um that's all just a bunch of baloney because um yeah we have latency anyways in the data centers that we train in and we have problems with data parallelism in the data data centers we train with and so federated learning has become quite a hot research topic and also a real production environment topic. Now many places use federated learning because of data locality regulation or other concerns about data locality that they don't want to move the data across continental borders or would be expensive or would be useless and some people use it for things like why would we even want to centralize the data in the first place. So when you start looking at um industry so factory work and this type of production you could see that um in even outside of learning for analytics or something we might want to push that to the device and I might only want to hear from you if this particular data pattern happens which shows me that a motor might be functioning improperly or something on the machine floor might be functioning improperly and only then or only when I run a federated query do I want that information in any type of centrally available manner. And so yeah, there's many different ways we can use it. But the cool thing is is that recent research by Google for an algorithm called DLCO showed that federated learning actually has quite a lot of advantages even when training massive models that cannot fit in one data center. So >> it's pretty cool.
That's super cool. And I'm I'm wondering what other what other concepts do you think or techniques do you think people should know about if they're going to start thinking about decentralized AI in particular?
>> I mean I would I would push to Joe to define federated governance because I think the big problem that we run into even with federated learning is does the data across our federated whoever population that we're trying to train with does the data match up? Is it reasonable? Do we have like a bunch of null values? Like what's the quality?
All of these types of things is like core part of how we could even try to begin to train um when we don't have a central actor, so to speak.
>> Yeah, it's interesting. I I'll be lazy and um actually see what uh how define it because she's the one that this is her fault. She started this. I'm kidding. um page >> and Joe also feel free to push any questions back to Katherine because you and I did set this up so we could we could >> you just kidding.
>> Um let me see here.
Yeah, I mean it it's I won't even bother. Um so you know thinking about what what she wrote and then what um you know what Winfrey's working on with his book on on federated governance it it's it's really a framework where again the if we're talking about like human um you know um centric governance for example right it means the the act of of data governance is distributed to people who touch data like I think right now the problem is with the centralized uh governance and Winfrey's argument is that it's simply it's not working right.
Uh there's this there's a bottleneck and it's a data governance work and with bottlenecks that creates loopholes. You can get around them yada yada yada. And so his whole approach is you know you people need to um you know people working on data products own the governance piece at the end of the day.
I think it's as pretty much simple as that. Now, now the federated computational governance part is one of those things where it takes into account basically what you were sort of describing which is like automated checks balances like is the data what I'm supposed to be getting right is it of standard form what is the shape of it what are the what are the um what does it contain what's the metadata and so forth so you know but that's but that's not that's that's still being built that's you know the thing that Jac's building over at next data right now right that doesn't exist well it does exist but it's early forms of it still right so But that was the missing piece. That's why data mesh that's one of the big reasons data mesh didn't take off at least in my opinion that and also I think people were approaching it wrong but it was um the federate uh computational governance because it's not just data within organizations right it's sharing data between companies and to a logical extent between people right so another person working on this kind of stuff is Martin Kleman if you don't know him who's in local first uh which is not a uh which is not a food co-op or a farmers market. It is a um it is local first uh computation, right? So if you can imagine Google Docs, but you have a local copy of it, you can edit it. Um other people can collaborate on the doc, right? But >> I think you mean it's not only a farmers market.
>> It's not only a farmers market. Yes. Uh yeah, you can um but you can get your locally grown data there. Um you know, you can export it and so forth, but at the end of the day, right, it's it's it's local. Um there's not an intermediary. So if you go through Google Docs and looking at notion right now, these are all companies that you know your data is it's there, right? So um we use we all use Google Docs, I'm sure, you know, just because it's convenient and sort of the default way that everyone shares documents now, right? What and your alternative is to go the other centralized route and share and email a word doc to somebody like a barbarian back in the you know last decade or so.
Um, my lawyer still emails me word docs.
I'm like, what? Like, what is wrong with you? I pay you good money. You shouldn't do this to me. Um, >> well, yeah, lawyers and tax dudes and I mean, sometimes they want you to come in for meetings.
>> Yeah, they do. Yeah, my we hired this CPA firm and then they wanted to meet.
I'm like, dude, I don't have time.
Seriously.
>> Yeah.
>> Yeah.
>> Hey, I'm in Germany, so yo.
>> Oh, yeah. Don't you still have the fax things or they got rid of that recently?
>> So, actually like the fax was the hack, you know, like if you fax you got seen sooner than email >> cuz there was essentially a policy that you answer the facts every day >> versus the email they will they will follow.
>> Yeah. So, they they changed the policy, I'm pretty sure. So, it used to be a hack and to get to the front of the line. When I when I lived and worked in Germany, I had a lot of interactions as Katherine um would know and recall from her own experience with the alander behen which is the alien office or something like what whatever it is, right? Like it's for like literally it's for the outsiders stuff. Um and I had to receive faxes from them and the only hack I could find was finding a service that would receive faxes and send them to me as emails.
>> Yep. Yep.
>> There you go.
>> Wow. actually holding Carell uh uh with her uh fight health insurance uh project. The crazy thing is I'm not sure if you're aware of like how this works, but to um have to file a dispute on your health insurance claims because that that is a thing in America. Um it's so you have to fax in your uh dispute. So what Holden has she set up a whole system which automates the uh the faxing of these things so you can just have her service do that for you which is bananas.
Um that's >> that is wild.
>> So before moving on, one more thing about faxes that has nothing to do with with this at all, but my my dad once received a fax from Hunter S. Thompson.
Um so my my dad was an academic and like cultural writer and and that type of stuff. Um and he reviewed Hunter S.
Thompson book very favorably. Um that had been I think panned elsewhere and he received like a gratitude facts essentially from Hunter S. Thompson Thompson. The the wild thing is remember that like glossy fax paper where like the ink would fade after a while. We we found it later like my dad had put in a particular place and it had disappeared.
It was a blank piece of glossy paper in in the end which we still have have somewhere but you know everything everything fades in the end including the the wonderful spirit of our Gonzo journalist. Um I so we've talked about um federated learning. I'm wondering are there any other things? So I'm I'm not I want to learn as well, Catherine. So I'm wondering is differential privacy and homorphic encryption and these types of things still still relevant to to think about today cuz I I just want to make sure people get a takeaway of the things they should start thinking about if they want to start working with decentralized data and AI.
I mean the the interesting thing is um obviously in federated learning many of the frameworks that we use today they exchange weights and depending on the type of data used for model these weights can leak information. And so when we start thinking about why or how we would use something like differential privacy or why or how we would use something like encryption so particularly for distributed protocols MPC or multi-party computation works quite well then this is one way that we would think of taking those updates potentially encrypting those updates sending them to a service that can work on the data encrypted add differentially private noise we're running that computation and then um decrypt essentially that could be what what I referred to in the book as secure and private aggregation what one cool thing so when I wrote the book now three years ago um that was still very research cutting edge that's now becoming or I'm starting to see more real projects is the idea that perhaps we don't need one centralized aggregator or even a myriad of aggregators that can work together there's more and more projects where I'm seeing the cool protocols in in kind of distributed compute protocols that have been used for a long time with decentralized idea to manage the workflows between multiple devices. And if we really think, okay, we're worried about privacy and security. Let's say we live in a country where we're we're worried about surveillance of our activities, where we're worried about um even communicating with one another might be something that we're concerned about.
And certainly the context of what we're communicating about we might be worried about or what we want to learn with like let's say we're exchanging messages or texts or we're learning with documents maybe even banned documents. These are all things that then we'd want to actually take a step back do real threat modeling and then eventually hopefully find protocols that do not have that one big aggregator because of course this um becomes the singular point of failure.
the singular point of attack. All of that is of course like if we're just if you're curious what I would ask you to first be curious before you step into federated is perhaps first be curious as Joe pointed out with like local first data. And this basically means setting up models on your computer using local models or even building yourself like a little load balancer which a lot of companies build in a production setting when they're like Joe said and I know like Hugo does use one model for this use another model for this use another model for this. So perhaps even formalizing that in a kind of piece of load balancer software I know there's many available. I've been meaning to do a comparison. Maybe this will bring me to be a comparison. And then yourself also for your own usage. This helps you decide who knows what. And when we have that load balancer, we could also start to think of privacy protections. So you could even, for example, have something there that sanitizes input. You could have something there that um rewrites input. So that's not exactly in your voice if you're sending it to an external model. You could have something there that writes it in the style of Hunter S. Thompson. Whatever it is that you want to do, you can start to play around with these ideas because at the end of the day, you're getting used to the types of things that you want to do locally. Then you can perhaps start fine-tuning local models. So personalizing them to yourself. And when you've reached that step, I think then you're ready to enter federated learning because then you've already been learning on device or on whatever compute you have available to you and stepping into how to do federated learning with friends becomes a a little bit easier once you've made those steps.
>> I love the idea of federated learning with with friends. We've got data seances. We've got federated learning with with friends. And I'm so glad that we're indexing more and more on on local models. Um I hate the term small large language models. So can we just call them language models? Um no one said that in this conversation just to be clear. Uh Joe mentioned OAM. There are many ways to to do this. Olama is um one of my favorites. Um and it's it's so easy to get up and running. Um even with you know when you don't necessarily have a lot of you know hardware on on your laptop. there are small models um you can play around with. I um I I love the small Gemma models for example and I'll link to this in the show notes but I've done a series of of workshops with uh Rivenkumar from DeepMind who works on all all the Gemma stuff about getting up and running uh locally with the Gemma models >> and even if you're uh on your mobile device too like duck.go as a browser they have uh the open source models so you can pick whatever one you want. Um duck.go go says they don't, you know, uh, store your, um, chats and you can delete them, so they say, so I believe them. Um, but it is a good way to get, I would say, like pretty innocuous questions answered if you're on the go.
Like I just use those, uh, cuz, um, you know, this pretty good. Brave even has their own large language model uh, which I think is or their own AI, which is kind of interesting. Um, it's okay, right? But um yeah, or you could just use Grock. That's that's really hinged and uh very sane. So um >> more like Tinder.
>> What's that?
>> I thought you said it was hinged like the dating app. So I said it's more like Tinder, but then I just realized maybe you were saying it's hinged.
>> Yeah. Yeah. Use use Crock for a dating app. That'd be pretty hilarious. You'll meet like just the craziest people, I'm sure. Um >> you'll definitely get what you're asking for and more, I think. Oh, I I was using corro because my all my friends were saying I can make really crazy images I can't make anywhere else. I'm like I don't want to know what you're talking about, but I'm going to try and see. So I um I won't say what I did, but it was uh but it it did work for the most part.
And I would say these are very unhinged things that could be banned in most countries. Um so yeah, but it will do it do it for you. And now you can't have woke AI in the US anymore. So now you're just gonna be able to create the craziest ever. That's fine. So, >> I was just going to say there there are more and more things arising like this, but I've got an app on my phone called um MLC chat, which allows me to download like uh smaller language models and and chat with them on my phone. Like it it burns battery like not as quickly as it as as it used to, but this is amazing, right? Like this is abs that I can run like a 3B model on my cell phone.
>> What kind of phone are you? Apple or you Android?
>> I'm I'm Apple and shamefully >> MLCD chat.
All right, I'm going to try this out here. Okay.
>> Um, awesome. Do So, Katherine, I'm interested. um doing this is is it an accurate characterization to say working with decentralized AI and C decentralized data um is is harder than not doing it.
>> I mean um I guess it depends how good the data is that you're working with in a centralized fashion. So if as Joe said like it's just a mess then it's just more of that except imagine that across many devices that you may or may not have access to >> that will make it better >> but I mean one thing we didn't touch on Joe I'd be curious to your opinions on this a lot of times when I when I'm meeting and talking with other like federated or distributed learning people um what they push or what I found to be some of the hardest parts is thinking through the identity and access problem which we'll leave to the side for now but then thinking through what I call like distributed data engineering problem which is basically let's say I already kind of assume that certain data might be available in whatever it is maybe it's a computer maybe it's a phone whatever it is maybe it's like a small chip on a on a on a like embedded device or something like this but I actually need to pre-process that in a meaningful way and to be able to ship that across potentially different compute environments and to have it run on each of them as the pre-processing step before the learning step can happen. And this is a non-trivial engineering problem right now. I'm wondering is there like a data engineering breakthrough that is like right around the corner that's going to solve this? I don't know of anything like that you know because that would entail having a consistency of a data model for example or transformation mechanism right and um schema registries data contracts put in a distributed fashion on edge I'm sure I'm sure people are doing this right but I I can't imagine this is something that you could just uh whip up um but you could vibe code it I guess and whip it up and then you'd be fine so uh maybe you could do it um That's >> I'm so glad you've finally come around to vibe coding Joe because, you know, we we've gone head to head before.
>> I I can't say I came around to it. I just sort of accepted that the world's crazy and it just is what it is. I'm I got other to worry about. So, I'm like, do you want to do that? If you want to blow up your production database like Jason Lemon did, go for it. Not my problem. Um I think I posted the other day, as an aside, too. I think the AI apocalypse is going to happen when uh you know just all these rogue agents uh you know and vibe coding tools just do like collectively do like drop table drop database and all corporate databases around the world simultaneously. Um uh >> we just got to get it trending via Grock, you know, and then the other ones will catch on.
>> They will. Yeah, that's the hive mind.
They all just align on that thing. And um but I mean actually that would be pretty devastating if you were to do that. So um but it's yeah it is interesting but you have the problem you suggest Kath that that's a that's a nextgen um type problem but I'm pretty sure right like that you know assuming decentralization continues or you know in fact though when you start thinking about you know the possibility of like humanoids and robots and all that that actually opens up the possibilities to a further extent for this decentralization because now you have data in the wild right and drones and all this other stuff and so I think right Now we've been so focused on um you know sort of this internal centralized database but as you have what Elon Mus is predicting what like you know he breaks a lot of stuff but like five humanoid robots per person right so that's you know potentially tens of billions of these things not including drones so yeah acting in a hive mind for example right assuming you have that with ambient agents and all this other stuff that is kind of out running in the wild yeah you're going to need a completely different way of handling data like stuff that doesn't exist right Wow.
Um, and so that's an interesting thought experiment. Like I've I've been telling people this like go ask AI like you know the fun fun thing to do right now is ask AI like if you were to design a data model an architecture from the ground up what would this look like? It does not exist right now. What it comes up with is very different than what we have. Um I can I can I had a chat with it last night cuz I was having a beer with my friend. Uh he was in town. What was this one? Um I can just read this to you if I can find it.
>> Yeah, talk us through some of the chat, man.
>> So it's interesting, right? So I asked it, um you know, I was in a movie about a month ago, 28 years later, good movie.
I was kind of bored one. So I asked if you were to reimagine data and AI infrastructure from first principles for AI agents, how would you design it, right? And it's like um so AI uh agent native data architecture real-time knowledge graph core storage event driven state management semantic data layers uh contextual compute orchestration multimodal fusion architecture hierarchical caching for reasoning layers and all the stuff that does like it's not here yet uh shared working memory um adaptive security and access control to Katherine's point um right so this would be intentbased security uh and so anyway way. It's um but yeah, but I asked I asked the that was the one I asked Claude and I asked the other AIS um the same question. They came up with very similar architectures.
So, um I'm surprised they aren't just building it behind the scenes right now.
They probably are. They're just like saying they're building like a dating app or something. I mean to be like overly boring about it. What it could also be is like if Olama introduced like a local data governance layer that was actually usable from a human perspective because we haven't had those in a a long time. And you could actually say okay here's all my email because you can download your email. Here's all my Google Drive because you can download your Google Drive. Um, can you make a workflow for me where you pull the document that somebody asks if I get an email incoming and can you only run it when my computer's on and give me the final okay check that for me to approve or whatever, right? Like we could solve it in very boring software oriented ways, but it's the matter of who wants to do that because that means that the data still does stay local, right?
>> Yeah, you're absolutely right. This could be done right now.
So yeah, I mean cloud code can just write it for you and just tell it what to do and it'll just do it and ironically then they have all your data and all your code and stuff but whatever it's fine.
Um, it's interesting.
But yeah, so I'm interested in just with the current focus on and hype around agents, which I, you know, like the ability of LLMs to be able to automate things I think is in or the ability of us to automate or be in the loop with LM that actually do do things I think is quite quite powerful.
I think they bring like vastly more significant privacy and and security concerns particularly when we're essentially asking of them to have access to everything from our chats to emails to, you know, Google Slides to whatever it is. Um, I'm wondering how how we can think about this, Katherine, particularly with respect to if we're using LLMs that are um more centralized than not.
>> I mean, I think like here's here's the problem is by design the way that these chat interfaces work is it's like a very intimate conversation, right? And I think that that isn't a random choice. I think that's like a very um >> quite interesting choice because even I was thinking about the other day like how come we don't have group chat available like where I could you know and there was we saw this with Slack and with numerous other things like you used to be able to invite bots in and this whatever but I think by design creating this kind of like private chat experience and even with the way that the inter most of the interfaces look they look very private and intimate and therefore for I don't know um how many people know of Joseph Fisenbomb's work but he wrote the first kind of like AI like chat therapy bot called Eliza you can look it up on Wikipedia Eliza and um he actually like threw away his development of the program because he found um one of his colleagues in the office like saying all these very personal private things to this bot and he was like found it abhorent because he was like it's just a computer like why are you sharing all this private thing and the person um I believe it might have been like one of the secretaries or other people that helped transcribe some of the the academic work he was working on said uh but it listened and I liked talking with it right and I think like very much that's what we see with um some of the also even agent design is this idea of this like helpful thing that can help us but from a privacy perspective that means I'm much more willing to share share more intimate information, maybe even information I wouldn't share with my friends or family because it does have this like idea that I'm talking with something that maybe can't tell on me, right? So, I think that's all by design. And I think the issue is going to be like how do we take the usefulness out of some of these things and maybe even some of the things that people are using it for like therapy or like you know to deal with whatever problems are happening in their life. How do we not strip that whatever it is that people need while still actually trying to point out that um you know sending all your deepest darkest secrets to a centralized AI chat is maybe not going to long term be the choice that you be willing to make and that's I think that's a difficult conversation that I feel like we probably should be having more often than we are.
Yeah, I guess you I was listening to a podcast on my uh hike this morning and they were talking about how people are using things like cloud code uh for non-coding exercises actually. So because cloud just looks at everything in your computer as a file. It doesn't really care. So you you can use it for all kinds of stuff actually like um use your imagination. So marketers are using it, sales people are using it like everyone you know I would say everyone but it's being used for non-coding use cases. Um, you know, somebody asked me why I don't like claude code, right? And I think it's it's not because I don't like the what it can do, but I feel like I'm giving it too much agency on my computer.
And that's, you know, I think that should concern you, especially if you buy things like Macs for privacy reasons, and you're going to let this third party tool just like run rampant potentially. Um, yeah, I don't know that that's uh smart, but um but you know, hey, you can write more code, so you can do that or do whatever. So, I guess that's, you know, but I think that that's that's a downside, right? I mean, I I I'm a big fan of things like um you know, you know, more like demos really, but you have operator and Claude's computer use and stuff and now, you know, more agents uh and whatnot, but like are you going to who here is going to let them just run free on your laptop right now?
You know, sight unseen, just do whatever.
>> I wouldn't.
>> Not at all. Yeah. And I never use cursor in YOLO mode. I know some people do, but I mean I I I give my agents read read access to, you know, sandbox versions of my local file system as well and but not not write access and if it wants to write, I make sure I'm in the loop there.
>> Mhm. That's just it. But the thing is all this data is being sent back to these companies like you know here's you know uh you know so that's that's just it like if you don't I mean so I think it's going to come down to like what how are you going to safely partition your your device so that it can act on certain things because right now I they say there's no you know there's there's no problem it'll be isolated I don't know that you can't prove that either right so that's the problem >> and finally just a quick shout out to Taon I don't I don't know if you all know continue.dev dev. Um, but it's an open- source AI code assistant that like think of cursor, but you can actually use it all with local models and I've linked to that in in Discord. You can use it with board and all all of that stuff, but they're they're doing super cool super cool work there.
>> Hopefully, there's more stuff like that, right? I mean, I think that, you know, these tools are powerful. I mean, I subscribe to cursor. I subscribe to co-pilot as well, and I think they're great. um you know and claude and everything else you use but it's like at the same time yeah I'm not not ready to give it control you know call me a boomer I don't care but it's uh you know computer use like I you know I remember um you know checking that out back over uh Christmas time and that was pretty cool actually like >> I had it in demo mode like I just gave it some instructions like find me you know awesome restaurants in this area and put it in a Google sheet and sorted by rating and it did it I was like cool that's pretty cool but I'm going to run it in the docker container there's no way I'm going to run this on my machine.
Like that would be insane.
>> Totally. And one thing I do appreciate about Anthropic is like they were very they were like, "Hey yo, like be very careful with this. This is incredibly dangerous, incredibly experimental.
Dockerize. Do not mess around people."
>> Yeah.
Yeah. Right. And so, but that's the next thing that's, you know, but we're talking about, you know, federated AI and now it's going to be federated agents, too. It's not just the chat bots anymore. It's, you know, ambient agents running around doing stuff behind the scenes. Hopefully they're doing the work you say. Hopefully they're not colluding and stealing your money and turning into Bitcoin. Um, you know, but that's that's the future we're in or the present really. I mean, it's it's crazy. I wrote the other day like it feels like every Philip K. Dick novel is like coming true simultaneously right now. And it's pretty fascinating just to >> Dude was a prophet.
>> He really was. I mean, took a lot of acid, but I think it may have helped him. So >> totally there's actually um one thing in the the book Do Androids Dream of Electric Sheep that um Bladeunner was based on that isn't in the film. There are a few key things but one is um there's a whole economy around robotic domestic animals. the world has no animals any anymore and they're like robot dogs that are um you know like I think incredibly expensive but they they start start to break down um and you know there are climate issues and all of these things and I think even these small details of his were so so precient in terms of even like robot dogs as luxury goods that like twitch and like have sparks fly out of their head. I mean wild wild stuff. Um, also a future where perhaps I mean there's an there's a potential future in which there's a whole part of the internet which isn't even human readable. It's only LLM and agent readable or said like MCP serves and clients or whatever the the next iteration of of these things are. Um, so I think like we'll there's a chance that we'll be more distanced from the infrastructure and how information is created. That's just it. Like, you know, I talk to friends, you know, especially in the knowledge graph community and and uh you know, the the library sciences.
They're like, well, you know, we we just need to have more knowledge. I'm like, I don't think I don't think the the AI isn't going to be speaking or language, right? They may be trained on it, but I think to your point there, my sense is there's going to be something else that they just invent because it's more convenient for them. Like, why be stuck with these human abstractions when you're a robot and you are capable of inventing new things? inventing new things. I mean, that's just it. So, it's, you know, >> yeah, >> to steal the argument for for building this type of internet and and these products, I I'm, as I'm sure you both appreciate, I'm all for um as much transparency as as possible and accountability. Um, and but um I I got like operator a while ago to go shoe shopping for me. Um, little known fact, I have size 17 feet at 51 European. You know this, Katherine. Buying shoes is incredibly difficult. It's also like I get triggered because as a kid, I'd go and try to buy shoes. Shoe salesman would say, "Why don't you just take the boxes home instead?" So, get having an having an agent that can like search the internet and dig in to different websites and and find shoes for me was actually very useful. Um, and it found a van store in Sydney which had a lot of shoes in in my size. But then I watched so it did a it did something for me that was very useful. Um, then I watched it do it and the operator spent like 15 minutes clicking around the internet to find these things. And I don't think that's the best interface for >> an LLM with tools to do it like to to do it. There must be better ways. But we still want need transparency and visibility and accountability into these systems.
>> Mhm. Yeah. Well, especially with the zeitgeist now that you know AI is going to take over every office function and and whatnot. It's like if you see a lot of interfaces that people work with in most corporate settings. I mean, they're a lot of clicking and they're horrible and they're there and you're probably not going to replace them unless you replace them all with something totally different. But yeah, so it's just going to be like toil and clicking and yeah, I'm sure I'm sure these agents are going to take up like heavy drinking and just like, you know, just really bad habits from this work. So >> they'll be going to LM therapists themselves and have like probably like rager holic agent like rehab groups >> and this is how they drop the database.
It's like screw this place. I'm I'm done. just like, >> "Hey, yeah, if I had all the problems of a modern agent, I'd you know, I'd shoot the schema."
>> Yeah.
>> Um, so I'm interested from from both of you, we've we've ascertained that decentralized data and AI isn't necessarily the easiest thing to to build and implement and maintain. So, I'm wondering what what are reasons and maybe and I can um ask you this first, Katherine, what what are reasons what situations will people find themselves in where it's actually pretty important to really consider these options?
I mean I I guess so when when I talk with people particularly let's say like creatives or other or even software people like us tech people like us and they think about what AI would be useful for them in the future they often talk about like truly thinking through an assistant role or even kind of like a subcreator role so to speak. And when I think about that, I mean, then you have a few options, right? You choose one model provider.
You link lock to them for life. You upload your entire portfolio. You trust that that information will remain yours.
Is you basically have the same closed garden platform problem we've had with every other social media platform, every other content platform, right? And um I think people know that's a losing battle.
And so or I hope we've gotten the message by now. I think even the influencers that like win that battle understand that is a very like the second the algorithm changes or the second the terms changes, you're the one that's that's disempowered in that information structure. And so what I ask of people is like do we want a future where your data and the things that you want AI whatever it is to do for you is actually guided by you and that you have actual choice and agency in how that's used and as you say and you know whether it helps you do shoe shopping or this or that or whatever I think then we must think through how decentralized works to some degree because um if we don't do that then we get locked into the same walled gardens that we've had essentially since the the mid to late 2000s. And I think that um that would be a real bummer I think not only for like maybe what AI could actually do because I think it could do many more useful things if we allow it to be personalized but also I think in terms of what does information mean in the world because if AI eventually helps us consume news or consume information that we may use to make decisions about our life then we want to at least steer that in some which way and we want to strive for diversity of thought in some which way and yeah we've seen that centralized algorithms and engagement metrics and these things don't help us do that very well so I would say from just uh your own personal health perhaps like for more cool things that we can think of in the future but also for societal health that it's useful for us to try to solve the hard parts I don't think they're unsolvable the hard parts of of federated or distributed Ed learning, but I think that we have to either decide we're incentivized enough as techies and like you said, you know, we build software that helps us do that or we have to put pressure that that allows that um monopolized environment to to shift or change.
Joe, I'm interested in if there's anything um you'd like to add.
>> No.
Um, >> some like there's a steamroller outside my house. I keep shaking my house, so I'm like like, you know, hopefully it doesn't uh Yeah. collapse my walls or something.
>> Yeah. I mean, I definitely hope that >> real world problems.
>> It doesn't.
>> Yeah. Real world problems. Yeah.
>> Um, so yeah, something I mean, we all want and and crave more agency in in our lives. So, let's not give those things up to agents. There there are things that agents can be good at um like doing doing certain things. Um but don't give up give up your agency. Um look at it's it'll be time to wrap up in in a minute, but I'm interested in this has been a wide-ranging conversation um not only about why decentralized data and why decentralized AI, but the types of things people can can think about and and implement um in their own lives and and at work. I'm wondering for technical people um who perhaps haven't had a lot of experience with these things um what are ways they can get started and of course I'll first say definitely check out Katherine's book practical data privacy which goes a lot deeper in into um a lot of these things but um what what other ways KJ um would you encourage people to get started today?
I mean, I think we've we've suggested a few pieces of of software. Maybe we can even make like a a list whenever this um is online of like different types of software that people can use. Maybe also in offline mode so that things are happening um without those those updates going back to a centralized point. So I think thinking through local first design for your own AI and then we've also mentioned some things that might be fun DIY projects which is adding layers.
I'm a big proponent lately of thinking through can we add a layer between the services that we use and what we're actually typing. So our prompts not only I think you I'd be curious your opinion here so that we can evaluate a variety of models against one another both local models but then also API based models because I think one of the things that is I'm already seeing consolidating and monopolizing is evaluation metrics across systems and keeping those for yourself either as a developer at work or also for your own personal projects keeping those yourself so you can actually evaluate I like this for this task, I like that for that task, or try out new versions. I think that that's really going to be the the one thing that can help keep your development somewhat independent.
And then moving from that, we've mentioned a few things that are unsolved problems or not yet fully solved problems in federated learning. So, if you're a big fan of Joe's work, maybe thinking through how do we ship data engineering workflows to distributed devices, right? If you're not already working on if somebody isn't already working on this problem and solved it, then write us. But thinking through like how how if you have a friend you trust or you have multiple computers at your house, which I'm sure you do, how might you take data from one of your home assistant devices, ship it to your local computer and then interact with that data via an LLM or whatever it is that you have, right? So, just think smaller first because you can actually see those devices or log in or if you have a friend you're willing to exchange data with, do that first. And and as you play with those problems, there's many easy reachable problems that just need to be solved by smart software and data people. And I think um that could be an awesome starting place. Um if you have a startup, then come say hi.
>> I love it. Um, and to there were several questions in there. Um, I I think a wonderful way to play around with and compare local models is uh using alarm with gradio um which is from from hugging face. It's a really lightweight um way to spin up um local apps that then you can actually share um web links for friends and colleagues. Um, and what I did as like I got a pretty beefy uh laptop recently and I can, you know, run, you know, six models in parallel on my laptop using OALMA and put one prompt in. And actually, I'll share um some code for how to do this using Gradio and see um all the responses coming through.
Um you mentioned modularity and and I I suppose LLM or AI pipelines earlier or hinted hinted towards it as opposed to just using one model in the middle using like maybe you have let's say you're building a multimodal rag system. Um let's say you have uh you want to get your images and text chunks in the same embedding space but you have different types of images. Um one being uh figures and charts, one being photos. split it and there are models that are far better at um uh captioning images and charts than photos. So splitting it in this way um and then seeing the output and then feeding that into your retrieval system um can can be really powerful. And the reason I also mention this is you can throw everything into one big infinite context like model but to iterate and then evaluate you actually have no visibility into the system at all. So having it into um all these small parts allows you to look at your end to end I suppose failure analysis uh and then introspect into which part of the pipeline is the biggest lever um to to pull there in order to improve your system. Um mentioning gradio I I think Joe I I may be wrong but did you mention or have we talked about that you um you you use AI studio Google's AI studio is that right?
>> I sometimes do. Yeah.
>> Yeah. I don't this is actually really weird. Um, can you I'm sharing my screen, but for the audio version, I'm going to just talk through what what's happening. Can you see my screen?
>> I could see two receipts.
>> Yeah, from Joe's Diner. This is This is something I was doing earlier today that had nothing to do with this, right? But, um, AI Studios actually got this quite it's got a lot of cool stuff, but um, I can put one prompt and whatever context I want in and see what the results of different models are. So this is Gemini 2.5 uh pros response and this is Gemma 3 27B's response and I could choose two different versions of Gemma there. So just quickly iterating in AI studio there's even I don't know whether it's in this view but you can even um oh yeah you click on this and it gives you the code so then you can actually paste it in uh to your own own system. So that's something I find quite useful as well.
Now, that isn't isn't local first, but it does allow you to play around with um >> open weight models as well. Um cool. So, yeah, that's something that I I thought may be useful to share there.
But you can spin this up really easily uh locally with um that type of interface using Gradio. And I'll share some code in in the show notes here as well.
Um, look, thank you both for your wisdom and your your way of being kind about, you know, a a lot of the hype in the space, but also helping us to manage expectations and see what what's really up, but also for coming and just sharing your time and having having a great chat.
>> It was great. It was, but yeah, thanks for thanks for having us. Um, yeah, we'll have to find a way to catch up in real life at some point, but uh, yes, fun times.
>> Um, any final words, Katherine, for people embarking upon their decentralized journey, data or AI or otherwise?
>> I don't know. Damn the man.
>> Love it. Well, um thank you all for for joining everyone who um watched live and everyone who's listening afterwards. Um and thanks for all the great uh chats in Discord as well. And see you all soon.
Winner.
Up Next

Federated AI Simulations with Flower: A 2025 Tutorial for Beginners
@flowerlabs
8.2K views•2024-12-11

Secure Multiparty Computation (MPC): Foundations & Challenges
@SimonsInstitute
7.3K views•2015-05-28

Bypassing Tor Censorship: Bridges and Pluggable Transport Guide
@Coding_ForEveryone
397 views•2024-06-11

Neural Networks Explained: Math, Layers, and Learning Fundamentals
@3blue1brown
21.9M views•2017-10-05
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Artificial Intelligence


























![AI 관제의 현실 해법, 엣지 AI Box [AI 관제 인사이트 ep.4, 라온피플, 토크아이티 웨비나]](https://i.ytimg.com/vi/XlubUda0uCo/maxresdefault.jpg)












