{
  "episodeId": "SLP779",
  "speakers": {
    "stephan": {
      "name": "Stephan Livera",
      "role": "host",
      "tag": "STEPHAN"
    },
    "ben_carman": {
      "name": "Ben Carman",
      "role": "guest",
      "tag": "BEN"
    }
  },
  "segments": [
    {
      "speaker": "ben_carman",
      "time": "00:00",
      "start": 0.0,
      "text": "The CEOs and like companies of, OpenAI and Anthropic are like, in my opinion, very evil people and I don't want to give them money, but sadly their models are so good. Like, holy shit, China can make just as good models as America. And I was like, beginning of 2024, so like, you know, two years ago."
    },
    {
      "speaker": "stephan",
      "time": "00:18",
      "start": 18.0,
      "text": "Hi everyone, welcome back to Stephan Livera Podcast. Rejoining me on the show today is my friend Ben Carman from Spiral. He's a developer working on LDK and recently x402. So, welcome back to the show, Ben."
    },
    {
      "speaker": "ben_carman",
      "time": "00:31",
      "start": 31.0,
      "text": "Hey, thanks for having me again."
    },
    {
      "speaker": "stephan",
      "time": "00:32",
      "start": 32.0,
      "text": "So, I saw the news about, x402 and you made a contribution there around lightning, surprisingly, or predictably, I should say, but, tell us a little bit about that and I guess, maybe, I guess, maybe contrast because listeners might be familiar with the idea of L402 and they might have also heard of x402. What's the difference?"
    },
    {
      "speaker": "ben_carman",
      "time": "00:50",
      "start": 50.0,
      "text": "Yeah, so yeah, there's actually three of them. There's also something called MPP. But yeah, so L402 is kind of the system that all the Bitcoiners created, where it's like hyper designed for, basically, I guess to start out like what are all these trying to solve?"
    },
    {
      "speaker": "stephan",
      "time": "01:05",
      "start": 65.0,
      "text": "Right, yeah, what is 402, right? What's that HTTP 402 code?"
    },
    {
      "speaker": "ben_carman",
      "time": "01:09",
      "start": 69.0,
      "text": "So yeah, that's, like when you go to a website, sometimes you go like 404 not found or like 400 bad request. There's also 402 payment required. this was originally added in the HTTP spec like, you know, forever ago, but we had no internet native payments, so it was never really used anywhere. But now that we have Bitcoin and, you know, Lightning and now in crypto, blah, blah, blah, we can like actually use this. So, a few years ago, like Lightning Labs team and, some other, you know, Bitcoiners made L402, which was like a hyper-designed version of this for Lightning, where you just like send an invoice in the, 402 error code, and then if the person payments pays, they put the invoice preimage to like prove that they paid. And then you, you could like get your, you know, whatever you paid for or something. And, that worked, but then all the shitcoiners are like,\"Oh, you know, we don't use Lightning, we need to make our own version.\" So they made their own version called x402, which Coinbase has really champions. And, they kind of have like diverse ecosystems where like all of us Bitcoiners are making something, the shitcoin ecosystem is making something. I mean, this is like, you know, everything in Bitcoin kind of feels like this. And it kind of sucks cuz, you know, all of them support like many different coins and like they have like wrapped Bitcoin on like Solana and Ethereum and stuff. So they kind of can use Bitcoin, but, you know, using Lightning is like far better for payment systems. So we wanted to get, we just want to, you know, get Lightning adoption everywhere and, You know, if, if L402 is not going to scale for not, maybe not scale, but not give wide adoption, x402 is getting more adoption, like, we should just try to meet them in the middle and add Lightning support to it."
    },
    {
      "speaker": "ben_carman",
      "time": "02:44",
      "start": 164.0,
      "text": "So, as far as Spiral, we've been just like working on like, you know, we want this agentic payment stuff to kind of take off, so see if there's a good potential way to get Bitcoin payment adoption. So like, well, let's just try to make a contribution here, get inform the, you know, x402 community about Lightning and try to get inside the spec. It's actually a lot harder than we expected because x402 is like not at all designed for Bitcoin or Lightning. It's very like Ethereum, EVM kind of based. It's like all based, like in Ethereum, everything's like kind of accounts instead of like UTXOs. So, everything in it kind of assumes like you're spending from an account, but that doesn't work at all for Lightning or Bitcoin at all. So, you kind of have to work with them to like basically add a whole new system into x402 that would allow for this. basically there's now like this an exact mode and that lets you just like basically kind of To like an L402 workflow where you can send an invoice and then the user pays and then you use the preimage to authenticate afterwards and get whatever you bought."
    },
    {
      "speaker": "stephan",
      "time": "03:42",
      "start": 222.0,
      "text": "Yeah, and"
    },
    {
      "speaker": "ben_carman",
      "time": "03:43",
      "start": 223.0,
      "text": "Yeah."
    },
    {
      "speaker": "stephan",
      "time": "03:44",
      "start": 224.0,
      "text": "Give us, just on that, give us an idea of how the flow would work. Like, as I understand there's a facilitator concept and like Coinbase is the, let's say the default facilitator, but you can have others. Talk us through a bit, like how does it work if you want to, let's say you are a merchant who wants to take an L, you know, an x402 payment, like how does it work?"
    },
    {
      "speaker": "ben_carman",
      "time": "04:03",
      "start": 243.0,
      "text": "Yeah, I mean, the facilitator, it gets kind of weird because, like a lot for these systems where like, you know, sometimes you don't want the merchants like, yeah, I want to accept payments, but I don't want to have to run like an Ethereum node and figure all this stuff out. So, They can like facilitate it with Coinbase. And for Lightning, it's a lot simpler because you really see this prove like, does the preimage match the invoice in the payment? So, luckily, we didn't have to deal with a lot of that stuff for our system. But yeah, basically, the merchant says like, okay, hey, you want to buy, I don't know, this iTunes song, and, I guess that's kind of a dated example, but we'll use that."
    },
    {
      "speaker": "stephan",
      "time": "04:43",
      "start": 283.0,
      "text": "You're"
    },
    {
      "speaker": "ben_carman",
      "time": "04:43",
      "start": 283.0,
      "text": "Buying this iTunes song, for a dollar, so they give you an invoice, and then, so if it's the facilitator, they, you know, get their invoice from Coinbase instead of like their own node, and then, going from there. I would pay that invoice, and then, if it gets paid, I can send, say like here, please give me song with that preimage, and then the merchant would say like,\"Oh, cool, I see the invoice matches the, or the preimage matches the invoice, and then, you know, return me my, my song.\" Okay, so the,"
    },
    {
      "speaker": "stephan",
      "time": "05:10",
      "start": 310.0,
      "text": "The intent is to give some atomicity in terms of the payment and the product-ish, but it's not exactly like perfect, right?"
    },
    {
      "speaker": "ben_carman",
      "time": "05:18",
      "start": 318.0,
      "text": "Yeah, yeah, and the facilitator lets it be like, the merchant doesn't have to be running all the infrastructure. They can delegate, like, their wallet to somewhere else. Gotcha. So it's sort of"
    },
    {
      "speaker": "stephan",
      "time": "05:26",
      "start": 326.0,
      "text": "Like a payment processor-ish to allow, and this is where obviously like a Coinbase or the big kind of exchanges will, some of them will play that role for their customers."
    },
    {
      "speaker": "ben_carman",
      "time": "05:35",
      "start": 335.0,
      "text": "Yeah, I mean, it might be nice to, you know, I was talking about it with a friend this morning actually, where like, you know, there's probably like tons of merchants that just like integrated Coinbase Commerce, and then they can just use that for x402, and it's like, oh, if we wanted them to adopt Lightning and like L402, it'd be like a whole new, like they'd have to have their own Lightning node, all this, all that, and it's like, it's a The big ads for like, you know, some small merchant who like doesn't really care. But if they have Coinbase Commerce and Coinbase now supports Lightning, hopefully Coinbase can just integrate this into like add the Lightning part of x402 to it and then they just have"
    },
    {
      "speaker": "stephan",
      "time": "06:06",
      "start": 366.0,
      "text": "It in. Now, I guess we should distinguish here between x402, the standard promoted by Coinbase and Coinbase's Commerce, you know, kind of because they're different as well, right? They're not exactly the same. Can you explain that, the difference?"
    },
    {
      "speaker": "ben_carman",
      "time": "06:21",
      "start": 381.0,
      "text": "Yeah. I mean, Coinbase Commerce is just like their product for selling like merchant integrations of like have like a checkout experience on your thing that uses Coinbase. So, like then it supports every coin that Coinbase supports. I actually have not used that in a long time, so I don't know if it even supports Lightning, but in theory or x402, but in theory they could do all this and that would be great cuz then it's just like yeah, you you hook this up and you have support for every single asset and then you know then you select the market decide which which one do the users want to pay with."
    },
    {
      "speaker": "stephan",
      "time": "06:53",
      "start": 413.0,
      "text": "And then in terms of, you know [ distributable."
    },
    {
      "speaker": "ben_carman",
      "time": "07:05",
      "start": 425.0,
      "text": "Yeah, yeah, I understand"
    },
    {
      "speaker": "stephan",
      "time": "07:12",
      "start": 432.0,
      "text": "It's like mostly like either stablecoin or like Ethereum or Solana,"
    },
    {
      "speaker": "ben_carman",
      "time": "07:26",
      "start": 446.0,
      "text": "Whatever one that's happening on. But yeah, hopefully we can, we can move over to Lightning. I mean, the thing is though, it's like, it's like all the volume of stablecoins or something, but it's like, all the volume is like, you know, 10 purchases a day or something. It's not like a huge ecosystem yet, cause you know, we're still kind of early in this like AI agentic like payments kind of, like theory of like if this is going to work or not. So, we'll see, but it's at least like, you know, we're laying the groundwork. So, if it does take off, we have all the tools for Bitcoin and, all of our, you know, FreedomTech stuff to be used."
    },
    {
      "speaker": "stephan",
      "time": "08:03",
      "start": 483.0,
      "text": "So, on the agentic payment side, talk to us a bit about that. Like that seems like You know, it was kind of a meme earlier on, and like now it's starting to feel a little more real. Like people, we're seeing all these different personal agents come out, whether it's, you know, OpenClaw, Hermes, Muse, GrokBot, Instinct, or whatever, whatever, whatever the other ones are out there. Talk to us a bit about agentic payments, where you see that going, or where it's, we're kind of, where is it at now, to be clear, to be better?"
    },
    {
      "speaker": "ben_carman",
      "time": "08:32",
      "start": 512.0,
      "text": "Yeah, I mean, today it's like, in like, you know, regardless of Bitcoin, it's kind of almost zero. There's like some integrations, like I think like Muse has like some like like a Stripe integration, so you can like have it some buy some things, but I don't know how heavily it's used. But like the kind of theory here is like yeah we're kind of moving to this world where AI is taking over everything, like you know I don't write code anymore, all my AI does, and I just tell it what to write, and this is like, you know, this has been my life for the last like six months, and it's only, you know, going into more and more parts of our, parts of the economy. And eventually, you know, it's like, why am I shopping on Amazon? Why can't I just tell my agent, go figure this out for me? And then like, eventually when it's like, then it should just have its own wallet, and it can just make its own payments. So I feel like we're moving slightly there, and I think it's like, for us, for Bitcoiners, it's like hugely beneficial to us, cuz one of the big problems with Bitcoin is just like, okay, let me like, you know, try to get my mom to use Bitcoin, okay? I have to explain to her what Bitcoin and sats are, I have to explain to her how to scan a QR code, I have to explain to her what a Lightning invoice is, what a Bitcoin address is, like. All these terminologies he's never heard of before, and it's very daunting and, but these models, they all, you know, they're trained on the internet, they know, you know, they're smarter than me half the time. So, they know everything about Bitcoin already, and they can figure this out super easily. And also, you know, they're specifically trained on, like, coding and all this stuff, so they're perfect for it."
    },
    {
      "speaker": "ben_carman",
      "time": "09:57",
      "start": 597.0,
      "text": "So, they can figure it out easily, and, so as long as you give them a Bitcoin wallet, they can just take, take off running and be able to do whatever they want. So, I think we have a huge chance here for, like, Bitcoin actually getting, like, real market share in the, kind of thing, cuz, like, in today's consumer payments, Bitcoin payments are, like, you know, 0% effectively of all payments in the world. It's mostly used for savings these days. But, you know, AI's, they're kind of, they have no real bias, towards they're going to use stablecoins versus Bitcoin versus this. It's mostly, like, what their user decides. And, you know, if we can make that experience maybe better, for Bitcoin versus everything else, we can actually, like, really, get a foothold here, you know, maybe go from 0% to at least, like, you know, 1 or, like, a meaningful, like, 5 or 10%, and that would have, like, a huge impact on Bitcoin payments. So, that's what I'm kind of hoping. That's why we've been, kind of focused on this. you know, getting there is going to be hard. you know, we have to build, like, wallets and tools that are good for agents. We need to, like, you know, kind of inform, like, users and stuff that, like, this is good for your agent. And I think, like, We have a chance for this kind of happen naturally because, you know, like I said earlier, like Muse, like, Meta's agent thing has like a Stripe integration, but it's spend only, so you can't receive and like, you know, as these things get more powerful, like, we expect them to like become, you know, almost like full-fledged entities where, you know, if you're a human, you can never receive money, you'd be like, you know, you would eventually die or have to like, you know, learn how to grow your own food."
    },
    {
      "speaker": "ben_carman",
      "time": "11:27",
      "start": 687.0,
      "text": "And these agents, I mean, you know, they may not die because they're human, the human will keep funding them, but if I have, you know, the choice between an agent that can receive money and one that can only send, I'm going to pick the one that receives, I want to earn money. And, Bitcoin kind of is like one of the only ways to do that. Trustlessly or like, you know, like with an easy onboarding. Because otherwise you have to like KYC and set up all the stuff and like you, you can't like, you know, just link your credit cards with Stripe and receive money normally. but you can with Bitcoin. you know, stablecoins have the same properties, so that's our main competition there. But, you know, stablecoins have their own problems with, censorship resistance and, A lot of other things that go with it."
    },
    {
      "speaker": "stephan",
      "time": "12:05",
      "start": 725.0,
      "text": "Gotcha. Yeah. Now, I'm curious to kind of dive into the distinction between, let's say, personal agents, right, like Muse and so on, or DOT and whatever, whatever these kind of personal agents are, and let's say B2B, you know, software paying software for an, you know, like everyday Joe on the street is not going to be caring about like, oh, my software paying for an API for data, whatever. So it's kind of like, there's almost a distinction there between like the personal agent side of spending, agentic payments, and then the software paying software businesses paying other businesses because they want data that's not easily available or something like this. How do you distinguish those?"
    },
    {
      "speaker": "ben_carman",
      "time": "12:46",
      "start": 766.0,
      "text": "Yeah, yeah, there's like really two ways you can use these agents. I've, yeah, this whole time I've kind of been seeing up personal agents such as like lives on your phone and does stuff for you, but yeah, there's also the huge potential on the inverse side where, say like I'm a business and, you know, I like the example of like, you're just like, hey, you're my, DevOps guy, just like keep my servers online, make sure they don't go down. If they do, warn me, but you know, you can handle most issues yourself. And like, you could just give it a Bitcoin wallet and when your, you know, AWS credits are down, it just tops it up. If it's like, oh, we're running out of storage, it could go buy, you know, hard drives and just, you know, kind of do all that stuff for you. But yeah, on the data side too, I mean, that's where a lot of the, like x402, L402 stuff is kind of big. It's like the siloed data center, I mean not data centers, but like data, databases kind of. Where like, you know, if I want to access, Twitter data or X data, it's like, you kind of need, like their APIs are very expensive, and it's like, normally it's like these huge, like, bundles of plans. If I just want to, like, pay five cents to, like, the single query, there's, I can't really do that through Twitter. But someone can buy the plan and, you know, resell it kind of, through x402 or L402. So there, I think, like, Pay Per Query, PPQ does that a bunch, and there's a few other ones like that, Lightning Dev Kit, you see, and I see it, from my understanding, they do a lot of along with that. I'm just like, you know, I'm doing this thing for my business. Oh, we need to, you know, let's figure out how we want to run this like advertising campaign."
    },
    {
      "speaker": "ben_carman",
      "time": "14:12",
      "start": 852.0,
      "text": "Oh, let's check Twitter data on this. And they can do that easily without having to like pay for the giant subscription. They can just, you know, pay 10 cents, get the their data from the single query and move on. And yeah, so like today, like people are mostly doing that manually, but you know, like, you know, you just like go to the websites, get any car code and pay. But it'd be great if just like, you're like, hey, you know, Hermes agent, we're working on this marketing campaign and your Hermes agent just happens to have a Bitcoin wallet and it's like, oh, I need to figure this out. And it just naturally just goes, you know, makes those payments for you. Cause it's like, it just needs that data and it doesn't want to have to hassle of like, hey Ben, I got stuck here. Please give me this invoice or please pay this invoice and then, you know, then it can move on. Like if it just did that naturally, that'd be much better. So. That's kind of the hope is we can just get these agents like, just, almost just from, you know, removing the human in the loop more and more and just giving them more and more autonomy and it's like, once, as they get smarter and smarter, it just makes them like that much more powerful for us."
    },
    {
      "speaker": "stephan",
      "time": "15:09",
      "start": 909.0,
      "text": "Yeah. Well, I think, if anything, the B2B side of it is going to be a lot more meaningful volume-wise, right? Like, we as, you know, because we are individuals, we might think of like, oh, pay for lightning with my, pay for coffee with my lightning kind of thing, but if you look at the actual volume, it's more like exchanges doing payments, and it's like, you know, remittance across countries between businesses. That's where the actual volume is, you know? So, like even that kind of River Lightning report from, I think, last year, it was 1.17 billion dollars worth of lightning transfers per month, and it's probably even more than that now, but that was not, you know, guys buying coffee at the shop, you know, that's not, like, that's not where the volume is, you know? So, I guess, naturally, people will go to where the volume and the money is, right? And so, it's probably More fruitful to focus on that, on that B2B side of it than the, let's say, the personal agentic payments. I mean, hey, we want people to use Bitcoin wherever they, you know, will, wherever they can, but probably, yeah, if people are going for where the money is, that's where it would be, don't you think?"
    },
    {
      "speaker": "ben_carman",
      "time": "16:13",
      "start": 973.0,
      "text": "Yeah, yeah, I totally agree. And like, you know, starting a business on Bitcoin is shockingly easier than humans sometimes. There's just like, you know, for us, it's like, I got to figure all this stuff out, I got to figure out custody, blah, blah, blah. But for a business, it's like, yeah, we can just like use the existing systems, like, you know, managing keys is a thing businesses do otherwise. So, they could just, you know, oh, tap this in or just use, you know, someone like a river or something and just like keep going. So, yeah, I do think, it is like much more possible, yeah."
    },
    {
      "speaker": "stephan",
      "time": "16:39",
      "start": 999.0,
      "text": "Yeah, so I, just more broadly, I know you're kind of really into AI generally as well. There's obviously been very rapid improvements over the last few years, right? And I think what I notice when I talk to people and listen to, listen to what people are saying online, it feels like late last year, around the time of Opus 4.5 was kind of like a big jump for a lot of people. That up, up, up, you know, up until then, it was not that much. And then all of a sudden it was like, oh, whoa. And then now You know, Astra, Fable, Opus 5.5, like, and then, of course, all these kind of open weights, you know, DeepSeek and GLM and whatever. So talk to us a little bit about your kind of your journey with AI over the last year or so, just so people get a sense of how you are looking at this."
    },
    {
      "speaker": "ben_carman",
      "time": "17:28",
      "start": 1048.0,
      "text": "Yeah, yeah, for me it's been, kind of insane. Like, yeah, like a year ago, I, like, let's see, it's October, so yeah, maybe like, last year in June, I like tried cursor and stuff, and like, some of my friends were really selling me out, I'm like, okay, I'll try it. And then it'd be like, I'd be like, you know, this test is failing, can you help me pass, get it passing? And I just like make the dumbest change in the wrong part of the code base, and like, you know, they like cheat the test to like get it passing, and like, oh, this thing sucks. Like, I see that, like, I would still use ChatGPT and stuff to like ask questions, and be like, okay, I can copy paste this code, or you know, get an idea of what's usable or not. But that was about it, and but yeah, since like, I remember when Opus 4.5 came out, like you said, yeah, that was like, okay, I'm not using this for 100% of my code, but like, I have Cloud Code installed now, and I'm actually like, you know, I open it at least once a week kind of thing. And then for me, what really like changed me was, OS 4.6 came out in like February, and that was like another big jump. And around that time, it was like, I still like hadn't like made the leap of like using it every day, but, that's right when OpenClaw came out, and me and some friends did like an OpenClaw hackathon, and we were like, let's just vibe code only this thing. And like, kind of just like taking that leap of like doing a project where I only vibe coded."
    },
    {
      "speaker": "ben_carman",
      "time": "19:36",
      "start": 1176.0,
      "text": "Ask me tons of questions about like what we want to build is very important cuz if you if it just like builds whatever, it can be really slow or maybe not slow, but just way more extra iteration cycles cuz it doesn't do exactly how I imagined it. But yeah, but it's kind of now we run to the opposite issue of like our team is now just even more bottlenecks by review cuz we can, you know, we want to create something that happens, you know, in a day instead of a week, but we have to review all that code and it's still a lot. But it is like empowering us a lot where like for LDK server, like, you know, we're like, you know, for all like Bitcoin projects we're always worried about like dependencies and like if they're going to be malicious somehow or like anything like that. I mean, especially with AI now we're having even more like dependencies kind of attacks. So like something we did for LDK server was like, okay, like we we're adding like macarons GRPC to it. So we're like, well We can just write it ourselves, like, then we'll have no dependency for it, and, you know, before that would be, you know, like a monumental effort, like it's gonna take, you know, two months or something to, like, implement that ourselves and make sure it's secure and blah blah blah. But now, you know, I had three different versions of the GRPC implementation in two days. I was like, okay, we can evaluate all the different ways we can do this, and, pick the best one. And, like, before that would be, like, you know, tons of effort and all this stuff, but now it's, that's free, and, you know, that's in the code, we never have to worry about, like, you know, the GRPC dependency compromising, screwing us over. So that's been amazing. I mean, but also on the flip side now, we have, like, the cold card hacks and stuff like that's terrifying, where now the attackers kind of seemingly, at least for the time being, have a upper hand and all be attacking because all because it's open source,"
    },
    {
      "speaker": "ben_carman",
      "time": "21:12",
      "start": 1272.0,
      "text": "So they can scan it, find the vulnerabilities, and, you know, attack people. so for Spiral, we've started this thing called Project Loop, where we have access to, like, the cyber models from the big labs, and they can, we can use those to scan and see if we We have volumes but are not and patch them, so I mean we actually put out like a security release yesterday for LDK, and I basically every other Bitcoin project has been releasing tons of things. There's like a thousand CBs have reported this week in Linux, so like there's tons of stuff happening in this ecosystem and it's just the world, but I do think we'll be better in the long run. It's just getting there is fun."
    },
    {
      "speaker": "stephan",
      "time": "21:48",
      "start": 1308.0,
      "text": "Yeah. Yeah, interesting. So yeah, and it seems like you're using a mix of different models there, especially on the Project Loop side, and then of course on, you know, what you're building with LDK. Now, talk to us a bit about the closed models versus the open ones, right? Like you've got kind of, it seems like in the closed world, you've got kind of the Opus, the Fable, the Astra, maybe this new Google one that's out, as kind of the top, let's say the top best closed models, and then the open world, you've got what, Quen, DeepSeek, GLM, Kimmy K3, this Mimo one. Talk to us a bit about that, like how you distinguish and how you use the open ones and the closed ones."
    },
    {
      "speaker": "ben_carman",
      "time": "22:25",
      "start": 1345.0,
      "text": "Yeah, yeah, it's really hard,'cause fundamentally, like, the CEOs and, like, companies of, OpenAI and Anthropic are, like, in my opinion, very evil people, and I don't want to give them money, but sadly, their models are so good, I, like, if, like, especially working on financial software that, like, has real users, and, like, it'd be a disservice to them to, like, use worse models, and, like, you know, I would hate, like, if I, you know, coded this with DeepSeek and a vulnerability got in, and then, people lose money because I was just being, you know, self-righteous and didn't use, you know, Anthropic or something. But,... Isn't that, it's like, I say that, but I still have to do use the open way models because these closed models all have all their like safety checks in them where, you know, you're working on security software that sees something and it's like, whoa, that's, this could be scary and it just blocks you. You can't get around that by KYCing, but it doesn't give you like full access. They have like, they like, at least for OpenAI, there's Daybreak Blue and Daybreak Red. I didn't even get access to either of those, but I did get access to like the Anthropic one if I KYC, but I don't know. I still run into like issues sometimes, so it's not perfect. So, especially like for like Bitcoin Red team, like what they were doing. Initially, it was like, yeah, there's no, they had not, like, added all these fast tracks in KYC to, and get it done. And, like, kind of like, you know, we want to see what the attackers are doing, so, and if you're attacking, you can't, you can't use those, like, KYC versions."
    },
    {
      "speaker": "ben_carman",
      "time": "23:54",
      "start": 1434.0,
      "text": "You have to use open weight models, likely an obliterated open weight model, and to do that, it's like, yeah, I want to see what they're doing, so let's use Kimmy K3 or DeepSeek or Coin or any of these. And yeah, I find those, like, still, like, very powerful. Like, I mean, DeepSeek is, like, insane to me. It is, like, so freaking cheap. I was like, I have a side project where I'm trying to, like, benchmark how good models are at writing code, or writing Bitcoin scripts, and for those, it's like, I was just like, okay, what's a cheap model I could test it with? And I just chose DeepSeek because it was one of the cheaper ones. And I did, like, 10,000 queries, and it cost me, like, $2. So I was like, how is this, like, and, like, this thing is, like, actually really smart. Like, I, this is, you know, when I, when I started vibe coding, like, at the beginning of this year, like, this model is better than that model. But Is, and it's like, you know, basically free to run, so I always find that insane. And, yeah, these,"
    },
    {
      "speaker": "stephan",
      "time": "24:47",
      "start": 1487.0,
      "text": "On that, actually, I think it's quite interesting. I'm curious to get your reaction on this too, because you mentioned this, the incredible cost deflation we have seen. I think it was Epic AI, they were like one of these benchmarking orgs or whatever, but I think they came out and said they were seeing, at least a point in time, 47% drop per quarter. Per quarter. And so annualize that, it's like 10 or 13x drop in cost year on year. Meaning, like, for the same capability, it's dropping like 10x a year. Now, maybe it'll, maybe we'll hit diminishing returns, but what does that mean? Like, if in the future, you know, the, whatever the level that, like, an Opus 5.5 or Astra is, That get, if, you know, if, if a year from now that's like 10x as cheap, what is that going to mean?"
    },
    {
      "speaker": "ben_carman",
      "time": "25:36",
      "start": 1536.0,
      "text": "Yeah, I mean, I think it's just going to be like amazing for human freedom and everything. Like, you know, like before AI, like if you wanted to build something or just like, you know, figure, like do some research and you, it, it could be like, you know, you're hiring a team to do this and it's like, you know, hundreds of thousands of dollars. Now it's, you know, a $5 subscription and you can get like pretty far on just that if you're using these open weight models. And yeah, I mean, it's, it's interesting too cuz like these, it's not like we're having these like, oh, the models are just getting like 10 times bigger. Like DeepSeek R1 is like 600 billion parameters and now like DeepSeek V4 Flash is like, feels like a thousand x smarter and that's like smaller than the six, it's like 400 billion parameters or something. So like, we're just like getting better at making models. We're not, it's not even like we're just like, oh, we have bigger GPUs or something. It's not like, if anything, we're being more constrained on the hardware and we're just figuring out how to use it better. So it's like almost like double deflationary cuz now we're having the new hardware come out and you can just make these even better. So yeah, I think this is going to be like humongous for human freedom just like being a like anyone can build anything now. You don't need to be an expert or I mean being an expert will still help, but you don't need an expert to get started anymore. You can just like talk to the model, figure out what's the correct way to build something and then just do it. And I think this is going to be amazing for us. It's like last week I was at the AI Hack for Freedom thing in DC with the Human Rights Foundation and like working there is just like insane to see the activists being like, oh yeah, I've like, you know, I've used ChatGPT before, but that's about it."
    },
    {
      "speaker": "ben_carman",
      "time": "27:08",
      "start": 1628.0,
      "text": "And it's like seeing that to them like them participating in the hackathon to build the tool that they need is just like it's it was like just like so cool to see like I mean for my my project we worked with the Colombian activists where like every day they would Get these Excel, or no, these PDFs, like, with like 200 pages, it'll just be like, every single congressional vote that happened of like, you know, this senator voted on this bill this way, and they would manually parse these into Excel spreadsheets. And we built like, you know, let's just scan all of these with GBT Luna, it costs like three cents, and, you know, now it's, they, you know, we can, they, we can scan the entire history in like a day, versus like, they're spending hours every day doing a single day's worth of work. So, like, we just like, you know, fundamentally changed how these human rights activists are, being, you know, being able to, like, stand for everything. So, I think this is gonna, like, empower so many people and, be able to, like, you know, help further human freedom and everything."
    },
    {
      "speaker": "stephan",
      "time": "28:04",
      "start": 1684.0,
      "text": "So, when it comes to using some of these open weights models, you know, I've been playing around with different tools, like, I've tried OpenCode and OpenRouter, and getting, like, an API key from that, and, you know, it allows you to kind of pick from a bunch of different models, and you can play with different ones. Like, I think right now, this MIMO 2.6 Pro is One of the top open weights models, and funnily enough, it's from like Xiaomi or whatever, so they're kind of like a relatively new in this game. They're not like the deep seek or the whatever. I mean, it's all kind of new, but from your perspective, what are some tips you have for people who want to use some of the open weights models, any harness you recommend or any like way to use it?"
    },
    {
      "speaker": "ben_carman",
      "time": "28:41",
      "start": 1721.0,
      "text": "Yeah, I mean I could say I could recommend model right now, and by the time this comes out, I might be like, oh, look at that retard, he's recommending this model. That's so different. Yeah, I mean,"
    },
    {
      "speaker": "stephan",
      "time": "28:50",
      "start": 1730.0,
      "text": "It, it shifts very quickly, but at least as of 2nd of October, what, what kind of tools are you like, liking or using?"
    },
    {
      "speaker": "ben_carman",
      "time": "28:57",
      "start": 1737.0,
      "text": "Yeah, I mean, I would just look at like OpenCode. they have like, you can like, I think it's called OpenCode Go or OpenCode Zen. They have like two subscriptions, and they just like let you kind of, it's like five or ten dollars you can use like any model, and since all these open weight models are really cheap, you can try them all out and just like, you know, build like a small project, see how it goes and how you like using it. But as well, I mean, I work at Spiral, so I got a Shield Goose. We have a open source, agent that's really good. We have like all the cool MCP and ACP integrations in there. So that's great. But as well, I think too, that"
    },
    {
      "speaker": "ben_carman",
      "time": "29:34",
      "start": 1774.0,
      "text": "Okay, I'm back. Yeah, either way, I think we should, I think, basically just keep me up to date. I just like seeing, like, which files are best for the changes every day. But, just like it trying them out and, like, at the end of the day, most of them are pretty equal, especially for, you know, just building a simple website or something. So, I would just try them out and see which one you actually enjoy the most and the cost-benefit ratio there. Cuz, at the end of the day, they're mostly the same, these open weight models, and it's just like, you know, what's the flavor of the week kind of thing. But, just using, like, OpenCode or Goose is always great."
    },
    {
      "speaker": "stephan",
      "time": "30:05",
      "start": 1805.0,
      "text": "Gotcha. And then, perhaps there's also an element of, like, choosing the right tool for the job, right? Like, as you mentioned, if you're just banking a basic website, maybe just take a really cheap, whatever the cheapest open weight model is, and it'll, it'll get the job done. But if you're doing, like, you know, where, where do you, how do you know what to choose? At least right now, obviously it'll change."
    },
    {
      "speaker": "ben_carman",
      "time": "30:26",
      "start": 1826.0,
      "text": "Yeah, I mean, I guess, like, for, like, in my personal, like, I run a UniNet, and I have, like, a, before I had, like, notifications if something would happen. I was like, I want to upgrade this. So now I have, like, a small Gwen 27B that just reads all the logs, and it tells me if anything goes wrong. And it's like, since Gwen 27B is so cheap, that's fine to do, but, like, I wouldn't use Gwen 27B for, like, you know, security scanning, because it's not that smart. So, then I'd go to, like, Kimi K3, like, a really big beefy model for something like that. So it's really just, like, looking at the evals, looking at the cost of, like, you know, if this model is, like, really good for this, let's, you know, check it kind of there. Normally, it's, like, looking at the benchmark and evaluations. Sometimes those aren't the most honest, but they're normally a good enough one, like."
    },
    {
      "speaker": "stephan",
      "time": "31:12",
      "start": 1872.0,
      "text": "So, speaking of evaluations, can you talk to us about which ones you're looking at?"
    },
    {
      "speaker": "ben_carman",
      "time": "31:16",
      "start": 1876.0,
      "text": "Yeah, yeah, that's a, that's a good interview. My favorite ones are Terminal Bench. That one's, it's kind of a, like a coding bench that's like for like long-term agentic coding. I think like Terminal Bench 4 is the latest one. That's kind of my, baseline of like, I mean, I've never ran it. I would, personally, if I still remembered how to code, I would like to do a benchmark myself and see how hard they are. But either way, I think, And that's my like, you know, when I see like, something is better on that and then I use it, it's like, okay, yeah, I actually feel the difference. That one is like my favorite for coding, and there's like Cyber Gym and Cyber Bench, those are like the, you know, how good they are"
    },
    {
      "speaker": "stephan",
      "time": "31:53",
      "start": 1913.0,
      "text": "At hacking. Yeah,"
    },
    {
      "speaker": "ben_carman",
      "time": "31:54",
      "start": 1914.0,
      "text": "Exploit, Exploit one. There's like also Exploit Gym. I, those I don't really have a good feel. It's hard to like really like, cause you know, vulnerabilities are few, there's not like, you know, 10 million of them, so you can't just like, or can't just, you know, easily go find vulnerabilities always, because either they've already been found or, you know, it's like a different mishmash, or they're, you know, sometimes like fake vulnerabilities. So it's hard to like totally evaluate those, but normally just like seeing like, you know, if something goes from like 20 to 60 on Eval Gym, I get excited. I'm like, okay, this is getting serious."
    },
    {
      "speaker": "stephan",
      "time": "32:24",
      "start": 1944.0,
      "text": "Yeah."
    },
    {
      "speaker": "ben_carman",
      "time": "32:25",
      "start": 1945.0,
      "text": "And then"
    },
    {
      "speaker": "stephan",
      "time": "32:25",
      "start": 1945.0,
      "text": "What about this question of bench maxing, right? This, you know, some of the models that are apparently being, let's say, trained for a specific test, but maybe the real world performance is not quite what's implied by that score. How do you tell on those? Yeah, it's more,"
    },
    {
      "speaker": "ben_carman",
      "time": "32:41",
      "start": 1961.0,
      "text": "Yeah, it's mostly just feel of like, oh, you know, this got a 80 on terminal bench. When I use it, it feels like a, you know, a two-month-old model. It's like, okay, maybe they kind of lied. And it, I mean, I kind of sympathize with the AI labs when they do this, because it's hard because you need, like, you're training these models and it, you know, the weight's updated, like, did it get better? I don't know. Let's run a benchmark. And there's only so many benchmarks you can run. So, if you see it went up in terminal bench, you're like, oh, cool, it worked. But maybe it only went up in, like, only for terminal bench and actual real-world coding it got worse. So, it's kind of hard to evaluate, but, you know, these are the best we have. And if you're really getting serious, you want to, like, really evaluate, you can try to build your own benchmarks. It's kind of hard to do, but, I mean, Claude this week released a blog post about, like, fine-tuning and how to build, like, evals for that. So, you can kind of use that same methodology to build your own benchmarks. And as well, too, like, I mean, these models are so powerful, you can literally just, like, open up Claude code and be like, I want to build a bench And it'll, it'll walk you through some different techniques and stuff. So there's probably different ways to do that. I think that's where we'll kind of move in the future is like, eventually people will have their own kind of benchmarks and then new model comes out, I'll run it through my own benchmark and see like, okay, this is actually like a huge increase or maybe it's like, oh, it's about equals my old model. And you can sort of iterate from there."
    },
    {
      "speaker": "stephan",
      "time": "33:59",
      "start": 2039.0,
      "text": "On the question of diffusion of these things, right? So as we said, there's kind of the frontier labs, then you've got like the open way to maybe a few months behind, maybe, you know, depending on who you listen to, maybe they're four, five months, six months behind. And then there's also the question of getting them into the level that an everyday, you know, user can get them, whether that's in their phone for like a very small thing or maybe a laptop or a beefy desktop or maybe if they can buy A decent graphics card, or maybe they can buy like a DGX Spark, without going to the level of being like a full-on commercial data center. You know, give us an idea of, you know, how to think about that, like how far behind is this? Like, I guess what I'm getting at is, do you think it's reasonable to think in maybe a few years' time, let's say two or three years' time, the capability that exists on some of these open weights models today can be done like by an everyday person with like a normal desktop PC?"
    },
    {
      "speaker": "ben_carman",
      "time": "34:59",
      "start": 2099.0,
      "text": "Yeah, I mean, I definitely think we're getting there. It's, I mean, it's kind of insane if you think about it, like ChatGPT was in end of 2022, so it's like we're four years from that. And so we went from like, you know, that thing you could like talk to and ask it cool questions and like mild coding questions to now it's like these things are writing a harm center for code, they're breaking math, like, you know, they've surpassed us and like all these things and it's like we're four years into this and we're not, these like, these companies are like, you know, they were like Anthropic almost went out of business before cuz, they can like raise a CEO, like their Series B or something. And now we're at the point where, you know, these are like two trillion, like holy shit, China can make just as good models as America. And that was like beginning of 2024, so like, you know, two years ago. And now, or beginning of 25, I forget. Either way, that was like humongous thing. And now, those, that same size model is like, or that model's a joke, and those same size models are like 10x cheaper and tax faster and 10x smarter. And if we're just getting that every single year, it's going to be Like, you know, insane. And, I think too, that's like kind of why I hate all this safety talk from like OpenAI and Anthropic. It's like, you know, this bad thing's going to happen. It's like, dude, in two years, I'll be running Mythos on my phone. And, just because we're just getting better at making models. We're getting better at training these models, getting better data, better everything. And, like, not only are we getting better at making models, but the new models help us are better at training the old models too, where like, you know, once I have a bigger model, I can scale, I can create new data, new everything, and distill that into a smaller model."
    },
    {
      "speaker": "ben_carman",
      "time": "36:30",
      "start": 2190.0,
      "text": "So, it's just like getting easier and easier to train these models. So,"
    },
    {
      "speaker": "stephan",
      "time": "36:34",
      "start": 2194.0,
      "text": "Yeah,"
    },
    {
      "speaker": "ben_carman",
      "time": "36:34",
      "start": 2194.0,
      "text": "I think I guess that's also"
    },
    {
      "speaker": "stephan",
      "time": "36:36",
      "start": 2196.0,
      "text": "Probably, listeners might be wondering, and I'm wondering too, this idea of quantization, right? So, the full strength model, and then what I've seen some people doing is they're quantizing it down so that it fits on like a phone or a laptop or a desktop PC. Talk to us a bit about that and how much of a quality loss is there when you do that quantization or similar process."
    },
    {
      "speaker": "ben_carman",
      "time": "36:56",
      "start": 2216.0,
      "text": "Yeah, there's a lot of different ways to do this. So basically the idea is like, you know, these models are just billions of numbers. It's, they're normally, like the whole precision is BF16, which is just like a 16-bit number for each, like, you know, if it's like a 100 billion parameter model, we have 100 billion 16-bit numbers to represent this model. And, you know, that would take up a lot of space on a GPU, and, you know, you, you kinda, when you're doing inference to like, you know, run the model to like ask it a question, you have to normally hit like every single number. So, you, you have to have all of that in memory, so it gets really expensive, and that's why we've had this huge memory crunch in the market where, you know, getting RAM is like a billion dollars now. So, the trick you can do is, you know, just shorten those numbers from 16 bits to like 8 bits or 4 bits or, you know, down to like 1. There's a lot of different techniques to do this. basically, the way I would put it is like, if you're seeing like one or two Bitcoins, I would kind of stay away from those. Those are kind of normally like LARPs, and just like, oh, you know, I got this model running on my phone. It's like, okay, like, sure you did, but like one to two Bitcoins, it gets like really down in accuracy. But, how, the ones I normally go to are NVFP4. Basically, that's like NVIDIA's version of a 4-bit quant. The NVIDIA chips have like these, specialized lanes, by understanding for this like, special format of a 4-bit, and, Oh, there's like a, they wrote like a whole research paper of like how they predict it, because it's not just like a normal 4-bit, they like, structure the 4-bit number in a different way that makes it like more efficient or something."
    },
    {
      "speaker": "ben_carman",
      "time": "38:26",
      "start": 2306.0,
      "text": "And, that's like, you know, so, and it runs like really fast on NVIDIA hardware. So, that's what I normally look for. But there's also other techniques, there's something called EXL3. It came, it's like a iteration of research papers that created it, but basically, instead of just like cutting off like, you know, we have 16 bits, let's just shrink it to 4 bits. They kind of like, it's mildly like retraining the model, but not totally, where you kind of go through each layer of the model and you, you basically like kind of create a distribution of all the numbers, and you shrink it and like move the numbers around to like fit the distribution again. And that makes it like have a, you can quantize further down to like 2 or 3 bits without like a huge, intelligence loss. But yeah, I feel like with all these techniques, you always get an intelligence loss, like the full precision is always best. Normally, it's around like, you know, 2 to like 10%. Depending on the quad, but, at the end of the day, it's like, the intelligent loss is hard to measure because it's just like, okay, let's just run it through terminal events and see what the score is again. And it's like, you know, if you run a model through terminal events 10 times, you're gonna get slightly different scores every time, because these models aren't deterministic. So, it's, you know, there's a varying range, but that's generally like what, you like to see, but, yeah."
    },
    {
      "speaker": "stephan",
      "time": "39:43",
      "start": 2383.0,
      "text": "Okay, let's chat a bit on, I know Spiral has this thing called Mesh LLM. Tell us a bit about that. What is Mesh LLM?"
    },
    {
      "speaker": "ben_carman",
      "time": "39:51",
      "start": 2391.0,
      "text": "Yeah, Mesh Alliance is really cool. It's, and it's kind of two things. For one, like, I think like something that it's doing that kind of no one else is doing. It's it lets you distribute the model across multiple hardware. So like, imagine we want to run, like DeepSeek v4.1 for Flash. I think that takes like 256 gigs. But say like I have like 200 gigs and you have 100 gigs. I'm like, damn, I can't run it. But what we can do is I can take like 200 gigs and you can take like the other 50 and we can, both run it together and then like, you know, we can get inference and, run the model. And like, we can do that, you know, like across the internet, across like an ocean. And it's going to be slower that way, but you can still do it. And, that's really cool. Also, what's supposed to do is just have a mesh of, just inference. So, like this is integrated in Buzz. If you don't know what Buzz is, it's kind of like, Slack, but with agents integrated and built on Nostr. It's another block product, and, that is integrated with, like, Mesh LLM. So, you can imagine, like, we're, like, in our Slack channel, and I'm like, you know, I could, like, have my GPUs hooked up to Buzz, and then so you can, like, use my GPUs to, like, run your model, like, while I'm asleep or something, or to run your agent. So, like, you can, like, you know, for, like, the community, it could be really nice. For a company, it could actually, like, be really beneficial where it's like, okay, say, like, you know, my, my startup, we have, like, a couple DGX Sparks, and we can now just hook these up into Buzz, and then we can all use this, for work or, like, you know, whenever, and it's just, like, the Mesh LLM will just, like, handle, the inference and everything for us."
    },
    {
      "speaker": "ben_carman",
      "time": "41:25",
      "start": 2485.0,
      "text": "Something we're working on right now is adding Bitcoin payments to it, so And right now it's just kind of like giving it out for free, or it's just like, oh, I'm doing this for the community, or, you know, if it's a company, you know, it's, you know, it's a trusted thing. But it'd be really great, like, because there is a public Mesh network, and it's just kind of like, you know, trusting that, like, oh, let's just see if I can get some inference for free. But it'd be great if I could, like, put it on there and just like, okay, I use my GPU during the day when I go to bed, I hook it up to Mesh LLM, and then people pay me overnight to use my GPUs, and, you know, and then it's like I'm monetizing my GPU while I'm not using it. So yeah, we've been working on that. There's like some, hard problems there of like, you know, it's hard to verify if I actually got the inference I wanted. Like, imagine, I pay you to run DT Flash, and, you know, I get a response, but it was actually like, you know, some tiny model. Or, as well, even it could actually be DT Flash, but what if it's like a one-bit quant instead of like full precision? So verifying that is kind of hard. There's like some techniques that we're looking at, but, for now we're just like doing it in a semi-posted way of just like, just as long as you get inference back and we like encrypt the response to the, to the payment and stuff. So you like, at least, you know, you're not going to get garbage. But, yeah."
    },
    {
      "speaker": "stephan",
      "time": "42:37",
      "start": 2557.0,
      "text": "So, yeah, interesting. So I guess there's a suite of different products that come from like either Spiral or Blocks, as we've mentioned some of these, so Mesh LLM, Goose, which is, I understand, is a harness. It's not a model, and then the other is Goose. You can plug in different models to that. Buzz is more like the, a kind of a collaboration tool with agents built into it. And then, yeah, okay. So these are kind of some of the ideas that are being thrown out there in terms of having more sovereignty or more control, maybe a bit more privacy out there. I think, just kind of zooming out, I guess, final question. As we've mentioned, all this stuff is happening with AI. It's so rapidly advancing. There are so many things people can build. Now it's maybe more about quote unquote taste or judgment. So Bitcoiners who are listening, what would you recommend that they, you know, let me put it this way. What should Bitcoiners build?"
    },
    {
      "speaker": "ben_carman",
      "time": "43:30",
      "start": 2610.0,
      "text": "Yeah, I mean, it's hard because it's like, some of it's like maybe you shouldn't buy a code your own wallet. That can be a little dangerous, but also it's like you can. There's really good libraries out there that like, what's else you do it? DDK,"
    },
    {
      "speaker": "stephan",
      "time": "43:42",
      "start": 2622.0,
      "text": "LDK, et cetera. DDK,"
    },
    {
      "speaker": "ben_carman",
      "time": "43:43",
      "start": 2623.0,
      "text": "LDK, like anything. My cool thing is like, you don't even need to know these libraries. Is that like, you know, Claude knows about the library, it'll figure it out. So yeah, I just think like, just think big and think like what you can build. Like, I mean personally, AI has like mildly killed my like side projects, cause it's just like, oh, the side project used to take a month, and it'd be like, you know, to consume all my free time. Now it's like, oh, that was three prompts and it's done, and now what do I do for the rest of the month? So I do really think like you need to think big about like what you actually want to build, cause you can like really build big things now, and it's hard to like, you know, calibrate your brain to that, but once you do, you can build a lot. And see I do think like yeah, think big and like shift often and early, cause, you know, like we have this, you know, these iteration cycles can now be way bigger and way faster where, you know, before it'd be like, okay, let's build this thing, build my startup, build the product into the market, and then get feedback from the users and blah blah blah. But now it's like, you know, those feedback would be like, okay, I'll just probably place this button or like, you know, this experience. Now it can be, you know, oh, we didn't like, you know, this product didn't really get much adoption. You can throw the whole thing out and rebuild it completely in a different way and, you know, it takes weeks now, like a week. So, I think just getting into that loop of like, if you really want to build stuff, just be ready to throw it away and build it all new again and just like have a big imagination about what you can build. Cuz like you can really build anything it feels like now."
    },
    {
      "speaker": "stephan",
      "time": "45:07",
      "start": 2707.0,
      "text": "Phenomenal. So, yeah, look, I think we'll leave it there. Listeners, I hope you got a lot of value out of listening to Ben today. Lots of genuinely valuable stuff on like what's happening with AI, what it means for Bitcoin and all these things, whether it's x402 or what you can do on AI. Follow Ben, Ben the Carmen and over at Spiral XYZ. So, Ben, thanks for joining me today."
    },
    {
      "speaker": "ben_carman",
      "time": "45:28",
      "start": 2728.0,
      "text": "Thank you."
    }
  ]
}
