Clojure's functional programming paradigm, with its emphasis on immutability and data-driven approaches, makes it particularly well-suited for data science and building computational tools for democracy, as demonstrated by the Polis platform which uses Clojure to process large-scale citizen feedback and identify consensus across diverse opinions.
Clojure for Data Science and Democracy: A Case Study
Added:thanks for coming everyone really excited to be here and it's an honor as always my talk is titled closure on the cyberpunk frontier of democracy and my name is Christopher small so this is a story in part about my very first closure application this is kind of how I got into closure and some things I'd like to touch on on that note are why I chose closure for this application what my experience with closure has been like over the course of its development and kind of finally I'd like to make a case for closure as a language for doing data science but this is also part of a story as sort of a new way of thinking about democracy it's a potentially new way of thinking about democracy and it's sort of a computational way of thinking about democracy so I hope that there will be some kind of interesting broader spectrum thought along those lines that I'll be able to share so no closure talk will be complete without the obligatory etymology slides so to get that out of the way the word democracy itself comes from the words demos and crotteaus which which translate roughly to the common people and to strength respectively so we can then kind of think about democracy is this idea of strength together right we're stronger together than we are apart but I think it's important to note that democratic behavior is nothing new it's actually ancient and we can observe Democratic like behavior kind of across the various species boundaries so non-human primates will exhibit certain kinds of democratic behavior wolves hunting will exhibit exhibit certain kinds of democratic behave in terms of deciding when they're going to go out to hunt bees exhibit Democratic behavior when they decide where to go build a new colony or where to go forage for for pollen and nectar but we often talk about democracy I visit having been invented or discovered by by the Greek and I think the reason for this is that really that they were the first to systematize democracy in sort of a you know a well a well formalized way and when I say systematize this is actually what I mean so this is a this is a diagram which represents the the structure of Athenian democracy from you know 500 BC and you'd be sort of forgiven for you know taking off your glasses and looking at this and think you were looking at a UML diagram right and what I think this highlights is that democracy really is a technology it's an information processing technology right we have people at the the center of this diagram here the the circle is the citizenry of of Athens and all these lines are going out in different directions represent the various governing bodies within Athenian democracy and and what's interesting is that a lot of these governing bodies were actually fed through sortition or through random sampling effectively right so they believed that wanting power anything in Athenian democracy was actually disqualifying for the position of power and so a lot of Athenian democracy was based on random lot random sampling and and so with what I think is really interesting about this is well this is an information processing and aggregation system it's actually built out of people right this is built out of institutions of people who who are who are collecting opinions and sentiment and sort of processing that into into decision-making apparatus but I think it's important to note that there are fundamental limitations to this and and the the Athenians sort of recognized this and thought that you really couldn't scale democracy beyond the scope of the city-state they they they're they're sort of framing of the problem was that if you were just bound by how many people you could get in a room physically or in a in a Coliseum in order to actually make these decisions together but aside from that that sort of kind of physical constraint there was also the constraint of different city-states having too many potentially different interests and and and the challenge of getting two different city-states with such different interest together to make to make decisions together would potentially be just too challenging and so at this sort of scale of things they thought that the best you could do is have sort of a loose Federation of city-states so this all sort of raises the question what would democracy look like if we invented it today if democracy is a technology at Balan by all these sort of constraints of computation then what would it look like today so I think it's helpful when asking this question to sort of take a step back and talk about cybernetics and in part this is sort of a hat tip to the title the talk right but the word and this is also our bonus etymology slide here so the word cybernetics comes from the Greek word for governance quite literally there's sort of some various variations of the word which translates something more along the lines of helmsman or ship steer which kind of gives you a sense for how they thought about about this word and what it meant but when we talk about cybernetics in the modern sense we can go back to its first English definition which was by Norbert Weiner and was defined as the scientific study of control and communication in the animal and the machine so the two words to really pay attention to here are control and communication right and one way of sort of boiling this down is that cybernetics at its core is really about any system that has feedback or signal which then is processed and influences behavior of the system so this is actually extremely broad and it's so broad is to almost become meaningless right so one can make the argument that is sort of generalized abstract nonsense as a result and indeed I think for a long time before I really got deeply into cybernetics and and really started kind of digging into the history that was sort of my impression of what cybernetics was right it's this word this sort of sounds like something out of science fiction and you know what is that really anyway but when you start to dig deeper you really see there's really a rich history here and cybernetics has influenced many many fields from artificial intelligence to robotics the Internet aka anything cyber star sociology ecology psychology I mean the list goes on and on if you go to the Wikipedia page for for cybernetics you'll see that there's just a wealth of different subjects that it has touched on and I think the value here is that in looking at things this kind of abstractly we start to see these system or sorry bigger bigger picture patterns in the way that systems behave and one of the one of the concepts that we can get from this which i think is is really potentially valuable is the idea of a signal problem the challenge of of signals getting enough information into a system to make good decisions so the first part of this here is that right it's difficult to gather and synthesize information from lots of people but it's also difficult to come to an agreement when different signals are pointing in different directions so these are sort of the two fundamental signal problems of democracy and governance this brings us to pull us so as I mentioned here in the slides but forgot to actually state polis comes from the Greek word for city state and so we name this sort of in harkening back to the the roots of democracy as we know it but really what the tool polis that we built is all about is scalable open-ended feedback so this is this is that feedback side of the the sort of cybernetic process here now one way to think about what this is really aimed to do is to look at social media social media has made it possible for you know millions billions of people to speak all at once and this is just phenomenal right this is you know if the Athenians had seen what we're able to do with Twitter or with with Facebook they would have you know they might they might have seen things very differently but the problem is that unless you're a data scientist you have no hope of being able to listen at this scale you know unless you're able to plug into the Twitter firehose and run a bunch of processing algorithms to sort of build some picture of what people are thinking and saying you have no hope of being able to understand that big picture and this is a real challenge right if you believe that the democracy depends on feedback of sort of information processes from the people this sort of lack of coherence is is a real challenge for functioning democracy so the question then is how can we use data science to make the opinion landscape transparent to everyone and so not just today two scientists who who can we can use these methods or the people who can afford to hire them but really to everyone how do we democratize these methods machine learning etc and and turn it into something that the entire society can benefit from what this actually looks like in practice is a participation experience where someone can ask an open-ended question so this might be something like what do you think about football injuries in and that is what do you think about brain injury in football or what do you think about gun control it could be anything right it could be very broad people respond with statements expressing how they think or feel about the subject and then they get the chance to vote on statements that have been submitted by others what we get from this now is based on all these votes we can put together this vote matrix where we have people along the rows and wear columns correspond to the comments in our statements in the conversation this by itself is not really particularly useful we haven't really learned anything yet but where we start to really glean some meaning is where we start doing dimensionality reduction okay so the problem with this giant vote matrix is that you have a bunch of points in this really high dimensional space and that's really sort of difficult to visualize we're we're pretty good at looking at two dimensional spaces and you know reasonably good at looking at three dimensional spaces although it can be a little challenging for computers sometimes but anything beyond that and we really start to have a breakdown in our ability to comprehend so the first step here is just to reduce the dimensionality of this data and the method that we use to do that is method called principal components analysis so if you're familiar if you're familiar with dimensionality reduction principal components analysis is a very sort of vanilla standard method for doing this and effectively what it does is you take this high dimensional space and you sort of rotate it around until you're able to capture the highest variance and you could sort of imagine there's an ax there's a analogy where you sort of think about a star field like a spiral galaxy or something and what you want to do is create a map of that galaxy well you wouldn't take the Milky Way galaxy and sort of flip it on its side like a dinner because you're gonna be missing a lot of the variance within within the galaxy if you do that what you want to do is sort of rotate it around until you really see the full spiral and now you've got a good map right this is exactly what PCA does it's just literally doing rotations in space until you come up with some lower dimensional representation that captures as much of the data as you can so because we're looking at opinion data here we can sort of think about this now it's like an opinion map or an opinion landscape and this sort of leads us to the next stage of our processing which is clustering so once you have an opinion landscape and a bunch of people represented within this landscape you can cluster them into opinion groups so each of these opinion groups then represents some sort of collection of opinions which tend to sort of go together but that's not the end of it where opinion groups really become interesting is where you start to dig in and figure out what's important to each opinion group and and this is this is really interesting in that it's it's one of the real challenges when we look socially at what happens on the web right when you're on Twitter and you see someone from one camp say something outrageous and and the peat and the folks on the other side sort of use it and hold it up and say aha look the other guys they're evil right look at this ridiculous thing that they said well that might not really be that representative of what that group really truly thinks in fields and really what disorder distinguishes them from quote-unquote the other side and this is this is this this this first step of just understanding what's actually important to each group is is really key to building a sense of meaning and understanding of of not only what your group feels and thinks but what the other group feels and thinks where were you sort of able to tone down the fringes but that's really just the first step because once you understand groups what you want to do is start to look for consensus between them and this is where the real magic lies in polis is that once you're able to break down the discussion into opinion groups you're able to look for where there's unexpected commonality those groups and that's often something that's very enlightening and sort of paves the way for pushing forward in decision-making in a way that's very challenging without doing this so what this actually looks like as far as the participation experience we have a data visualization which shows each of these little sort of circles here is someone's position in the opinion landscape and these are you know Twitter or Facebook profile images and you can explore this you can click around if you click on a group you can see what's important to that group and they're also some UI components which aren't shown here which lets you see those consensus comments and again I really just want to highlight that the real secret sauce here is consensus because this is what brings coherence to to a discussion and and what lets you see that oftentimes there there are pieces of their ideas or thoughts that that actually bring us together more than our sort of tribal social instincts would would have us believe so as sort of a case in point of where this has been used I'm gonna tell you a little story about Taiwan so in 2014 Taiwan stuck a trade struck a trade deal with China and it just had a bismil national support no one really liked this thing and part of the problem was that it had not gone through sort of the standard review process where people could sort of comment etc and part of the issue here is the Taiwan sees itself as an independent nation while China uses that use it as sort of a rogue province and this worried people that this was sort of a step towards you know a slippery slope of being sort of reintegrated into China which was really concerned a lot of people in the country so they saw this as a failure of democracy and what resulted from this was a set of August Occupy style protests where they occupied Parliament building and they occupied the streets outside and all of these old school facilitators were able to come up and in the streets with with giant poster boards and sticky notes playing these old school facilitation methods sort of from the ground up building building ideas about how they could resolve these issues and they were actually able to work out a more agreeable trade deal and get the government to agree to it so this was sort of phenomenally successful of all the sort of Occupy movements that have happened this this was I think one of the few that's really had some sort of tangible positive outcomes and once once they got there everyone went home and everything was everything was great but when the key sort of points here was that in this whole process this the civic tech group called gum zero g0v here came to came to favor with the population for the role that he played in these protests they were wiring Ethernet out into the streets and like setting up makeshift Wi-Fi antennas with tin cans and stuff and making sure that everything was transparent they were recording everything and broadcasting it and what this meant was that no one could say oh they're going and tearing things up or they're you know misbehaving everything was transparent and so no-one could sort of cast a false narrative about what was going on in these protests unfortunately though the government was sort of suffering a crisis of confidence in the wake of of this event and what had happened and in response to this the then Minister of digital Minister Taiwan Jacqueline's I went to a Cobb stereo hackathon him you'd imagine like some notable like repetition coming to like a closer hackathon and asking people to help out this is sort of paints the picture right she came to a gov gov zero hackathon and challenged them to build a platform which the entire nation could participate in in order to deliberate issues of national importance and the response to this was a tool called and not really a tool more of a um a platform or suite of tools called V Taiwan and there's a number of tools that sort of fit into V Taiwan and it's sort of a open-ended process that depending on the details can go in different ways and use different different different methods but the part of V Taiwan which certif alleles this goal of having you know having a platform that the our nation can deliberate within is is the polis tool that we built and so this sort of sits sits at the the early phase of of the V Taiwan process where they want to get large-scale open data and feedback from large groups of people so this has been really successful they've used this to regulate to figure out how to regulate uber and airbnb which of course is really interesting because this has been a real challenge for a lot of countries figuring out what to do with this sort of new technology and it also helps them break through issues on which they've been in gridlock for four years one particular case here is this issue over online alcohol sales in Taiwan is something that the nation had been debating and been in gridlock over for I think four years four or five years and within a few months using this sort of open process they were able to actually build from the ground up a set of a set of positions which helped sort of stake out you know a direction that they could the key building in policy and what I think this shows is that when you sort of take out all of the political maneuverings and backroom dealings and horse-trading etc there are actually pieces of of the kind of collective conscious that can bring different groups different sides of an issue together and and really what happened here was there were specific statements from either side that sort of gave a little room and said well maybe if there was this kind of regulation in place or maybe if there was this sort of need being addressed we'd feel more comfortable about it and that was what sort of helped them push through to to a proposal which had really broad support and again this was something that the politicians and the traditional you know representative establishment was not able to break through for a number reason so I could talk about that later if it folks have questions but but this is really cool right so closure and polls are being used now to make decisions of national importance in this country Taiwan so if you're interested in reading more about this I'd recommend this there's a number of article has been published on this now there's we have a wired article which is great and some stuff in the MIT tech review but really the first english-language publication about all this right because this happened over on the other side of the planet where big in Taiwan and no one here knew about it right unless you spoke Cantonese this was like this was this was this was something that no one knew about but Liz berry it's really wonderful individual who I think actually is from North Carolina here and she found out about this and I she went over there and was interviewing people and and wrote this English article which which then has been kind of a keystone for us in terms of other people in the english-speaking world finding out about it and since then you know we've started to work in Canada and Singapore and we're doing some stuff with local local newspapers in the United States and and also within companies and are really you know making some great headway so if you're interested in reading more I'd recommend checking that out but so far I haven't really talked about closure yet and you're all maybe wondering what am I doing here is it supposed to be closure conference so so I'm gonna switch tracks here a little bit and and talk about that so the initial implementation of the math engine behind polis was written just as a little R script right this very simple little thing that we ran in a loop and this was fine as a proof-of-concept it worked but we knew that it wouldn't be scalable as the system grew as we wanted to do more kind of in-depth analysis you know we knew that we'd read we'd have a harder time kind of working with the data in this paradigm so a couple years before this point I had started getting into functional programming and just kind of as a side hobby thing it was sort of felt good for the brain like I was learning a lot and in particular I started started using Haskell and I really loved it it sort of tickled the the math part of my brain and I was really interested in thinking about you know could be this is something we could use for for building the math engine behind this but I wasn't really sure if this would be a good fit for doing this kind of work in particular I wasn't sure if the you know we'd have performance issues or you know or are we gonna be yet looked at crazy and not be able to get funding or that's sort of that sort of thing but one of the things which sort of tipped tip the tip the scales towards me sort of more seriously considering this direction was an article that I read called the downfall of imperative programming and the the line here which stuck out to me was concurrency and parallelism are the killer app for functional programming so this was a really compelling idea the basics of which were just that immutable data eliminates this entire class of issues associated with concurrency and parallelism and we saw that this would potentially be a win for an application that's going to be dealing with a lot of data and that we want to be sort of responding in real time and and and so we started to sort of give it more serious consideration we also thought about a number of other functional programming languages Scala oh camel but for a number of reasons we sort of converged on closure and I think one of the big ones was that again because we're sort of thinking about this as you know how can we utilize concurrency and parallelism we recognize that closure was built right with one of sort of its primary motivations as a language for tackling the challenges of concurrency and parallelism and so that really appealed to us but another feature is that we saw that being hosted on the JVM we'd have the reach of deployment of the JVM and we'd have the richness of all the libraries and sort of ecosystem pieces that that we have available in the JVM and in particular right the JVM had all of these great like big data processing libraries Hadoop and spark and and all the sort of stuff you know we really saw that that was a strength of the JVM ecosystem and that using closures Interop we'd be able to take advantage of that but really one of the most important pieces was just that it was fun learning closure was really fun in a way that I mean I again I really liked Haskell but there was sort of the sense of austerity with it right where you know you get in and you sort of have to fight with the compiler until it compiles and usually works but but there's something just really fun about working with closure had the feel of of Ruby or Python but had all this great functional stuff too and that really appealed to me and especially when you're building a project as a side project passion project you know when when you've been spending eight hours working at your day job and come home and want to spend a couple hours on a side project you actually want to look forward to it right you don't want to be sort of dreading cranking out more code in a language that's that's um that's not you're not not not really hitting it for you on that side so the question is how did this work out that's the theory right great obviously I wouldn't be here talking about this if if this hadn't gone well but to dig into this a little more I think one thing that I didn't see as clearly heading into the closure ecosystem that really quickly began to I began to appreciate more deeply and which I think still I start to see more and more deeply the more I use closure is this idea of data-driven programming and I think this is something that really sets closure apart in the functional programming world is this sort of focus on programming with data and one sort of look at this is that it's sort of the Lisp philosophy taken to the next level right it's a more sort of dynamic take on the Lisp philosophy there's also libraries like prismatic plumbing this is sort of a little example here right that was there were really useful to us and allowed us to specify the structure of a computation separate from its execution so we could write the the express the computation is this data structure which we could then actually modify and mutate not mutate but you know compose and and sort of build up in different ways and then specify how it's going to be executed whether it was going to be lazily executed or greedily executed or executed in parallel or executed in parallel with profiling etc and that was really nice and and and felt very powerful but kind of at a really base fundamental level I think what what stood out was that data science really naturally benefits from this kind of dynamic data centrism you know if you're working with data and that's kind of the core thing that you're doing I'm having a language which is very like data driven felt really natural there's also kind of in tandem with this this kind of growing recognition of deep thought in the design closure where you just get the sense as you sort of dig in that everything had really been carefully thought out and all these pieces had really carefully been put together towards something that would be a really high leverage and composable tool and and this word leverage here's one that really sticks out to me right like this when I when I think about closure I think about leverage and composability is a big one too I think the composability is what gives us this leverage but but the leverage is actually the thing that we care about as practitioners right we want to be able to with a very small amount of force accomplish a lot and I think that's one of the things that that this collectively gives us so I'd be remiss if I didn't also express some of the challenges that I feel like we had over the course of the project again Craig this is kind of a retrospective and sort of looking back and thinking about what was what went well and what was what we had to learn from I think the biggest of this was just that you know anytime you're working with a new language or a new sort of system there's gonna be sort of challenges and you basically you build some stuff you learn a lot and then you look back and realize how bad that you build stuff and how you might have done things differently and you know I think I don't think this is necessarily anything specifically on closure I think this is more just you know when you learn a new language is this kind of something that's bound to happen you're you're learning these new patterns and and methods for doing things and this is also kind of hedge if anyone goes to look at my old code that you know some of it might be six years old now so don't don't judge it too harshly but some other issues kind of come down to specific libraries or frameworks that we thought would be really useful that we'd be able to leverage that turned out not to be as as great as as we found so one of these was in Cantor which is a really ambitious project that sought to sort of implement a lot of the kind of stats and data visualization functionality that I was used to working with are but unfortunately at a certain point its development it kind of slowed and stalled and it wasn't you know they weren't releasing new updates and we were having problems with dependencies and such there were also broken routines like the PCA routine for a long time was broken took a while to get that fixed I think it's they've maybe finally fixed it now but but also their websites still not up to date and so you know it's really hard to kind of figure out what's going on where but another issue here was just the limited graphing they did have some graphing routines but they were sort of very primitive and you you know in in the I'll talk more about this later but in the our world there's a really wonderful library called ggplot2 which which I'd grown accustomed to and it's just very powerful and the kind of stuff you can do with it it's very extensive and I really missed that in in this ecosystem so there was some pain points along the way with that and eventually moved kind of moved off that and sort of you know we implemented our own a PCA routines and and you know had to had to Co go some different directions there another thing that we thought was going to be really useful to us that turned out to be kind of more trouble than it was worth was storm it was actually one of the initial reasons that we sort of went in this direction we saw Oh storms this really cool like you know big data stream processing library that's actually written in closure or at least initially was and then kind of more in the JavaScript direction and her site Java Java direction and unfortunately here we just had all these issues with the head of time compiling and and dependency issues and as the development on the project slowed as it became an Apache contribution project we found that the docs were out of day and they're focusing more on the java api than the closure api so it's really hard to figure out how you were actually supposed to do things now when when when versions would update and so ultimately we just dropped this we realized that we didn't really need it yet and it was just kind of adding more trouble than it was worth so so that was that was a bit of learning along the way another thing here was that and again kind of this is sort of ahead of time compile issue and the way the things were running on the Heroku platform which is what we used initially to deploy this we found that you know just starting up database connections kind of in line in you know a def form had problems when we were deploying and so we look towards Stewart see our components to kind of separate the the structure of the computations from the execution of the startup routine and so you know I actually really like the design of Stewart Sierra's system components I think it's a simple solution to the problem and I really like the the idea of the system as a value but unfortunately the the the process of sort of taking code that had already been written and porting it over was really painful and took a very long time and you know it's time that we could have spent working on features and other other development and another sort of side of this is that it really makes interactive repple development a lot more painful instead of being able to just kind of start hacking away in a file and executing some forms using whatever kind of in in editor repple connection you have in my case I used vim in vim fireplace we had to sort of have these little workbench chunks of code where you'd start the system and then like get the system component you need and then pass that into the function that you wanted to test out and it was just a lot less sort of fluid of a development experience than just being able to do this stuff in whatever namespace file you were sort of developing and tinkering with at the time so that's those are some kind of the pain points here but unfortunately there's a silver lining again I probably wouldn't be here if this was all pain points so to move along with that there are now some alternatives to the sort of traditional Stewart Sierra system components one of these is Mount which is sort of a really elegant approach and has really I think very friendly API and solves a lot of same problems you do miss out on some of the value of the the system is a value thing but ultimately in other applications I've worked on now since then it's it's been a really much more of a pleasure to work with and so that's one thing to consider there's also a library called integral which has a very kind of different take on this problem and it's really interesting so depending on what kind of what kind of tool you're building whether it's a library or an application you know these are some other alternatives that you can consider additionally since a lot of this stuff has actually been pretty recent but there's been sort of this growing trend if I think other other folks getting into closure for data science and machine learning and such and we now have a number of neural net libraries and actually I see Karen Meyer out here she's been working on closure timex net which is which is a closure sort of API for the MX net library there's also a cortex for Mike Anderson flame from Aria and there's a closure their closure bindings for the deep learning Forge a library I'm so again know a lot of some of this is just taking advantage of the the JVM ecosystem right but the fact that there's so much attention being paid this is I think it's a really positive sign that that the closure ecosystem and community are kind of moving in this direction there's some other really great libraries that have come up recently kik see stats has some really great work that kind of fills in some of the niches that that we saw in canter potentially having having solved for us there's also the Anglican probabilistic programming languages which is really cool especially if you're into Bayesian statistics finally the core matrix library was kind of what we ended up leveraging when we moved away from encounter we just went kind of directly for this quarter matrix library and implementing these these various routines ourselves and one of the really nice things about this is that the court out matrix implementation is such that a lot of the machinery is to find around protocols and this is you know another one of these design features of closure which you know I think afford just a lot of really high leverage in that you can actually write all this kind of general matrix code and swap out the underlying implementations so depending on what kind of work you're doing and whether you need more performance or whatever it can in a lot of cases just be one line of code change to switch from one implementation to another and things just kind of seamlessly work now that even said there's some caveats with this right and I'll point down here to Neanderthal because it's it's a more sort of low-level approach to matrix work in enclosure and it's more sort of closely tied to some of the the underlying like glass and then your algebra pack routines and so if you're looking for something kind of more bare metal performance Neanderthal is a good approach for that there's actually one again though one of the great things about quite a matrix is because it's so abstract you could actually and people have been working on this you can actually implement the core time matrix protocol protocols on top of Neanderthal and there's some potentially issues with this if the protocols it's possible for the protocols to sort of be leaky as far as performance considerations go so you know you kind of have to be a little bit careful here if you're going the cordon matrix route but but it's nice to know that when you need it there's there's some folks here who are working on some libraries for doing more kind of low-level high performance work with closure another thing that I think was has been really a surprise to me that I didn't really foresee coming was that the closure script ecosystem has been just amazing and you know I had some I had some initial skepticism here I wasn't sure that it would really be possible to build a sort of a mock of closure on such a different runtime as JavaScript and have it actually be something that didn't feel like a pain in the butt but once I tried it I was really blown away and it was amazing to me that most of the kind of core language features were exactly the same and so now with with clj C we can write closure code that either runs on the JVM or closure script and you know with the ability to change just a little bits that that are different here and there and this has been really wonderful there's some really excellent tooling here I'm particularly you know use fig wheel reagent or reframe but these these libraries collectively really knock the socks off of working with vanilla react applications and I've done I've done a bit of this now both in in vanilla reactant with reagent and reframe and and I can say that you know things just move so much more smoothly in the closure script world so this has been this has been a real joy and a pleasure another sort of side of this is data visualization so I mentioned that in cancer had sort of let me down on the data visualization side but recently I've sort of taking a second look at Vega and Vega light Vega has actually been around for a while I think Vega white has too but just in the last year's so Vega light sort of released a new version that solves a lot of these old issues with with their initial implementation their old architecture and and I've really been blown away with it so it's based on these are JavaScript libraries and so they run in the browser but they're based on the same grammar of graphics idea that the ggplot library I mentioned is based on and the idea is you have this grammar of graphics which tells you at a very high level how do you translate properties of the data that you're working with into aesthetic attributes of your data visualization so I want a color by this very or this property I want to make the shape correspond to this property etc it's really easy to do this at the high level and not be monkeying around with a bunch of imperative code but it also comes with a grammar of interaction which lets you do really powerful things like you know select a bunch of points in a brush selection and then have that update some other part of the visualization and part of this is that has a complete dataflow specification which which facilitates this which is really really amazing but what's really cool is that all of this is specified just as simple JSON data right the the it's just a pure data json api which is really amazing because what this means is you can take that json and you can pass it from one process to another so i can write vague and fake a light code in closure or closure script as eden and translate it over to json and send that over a rep or a socket to to to vague or vague a light in the browser and have it render that data and this is really powerful i think and really goes conceptually hand-in-hand with closures data-driven philosophy so I'm running a little short here so I'm going to skip through a little bit of this but Vega is really super powerful and customizable but it's not super well suited for day-to-day use Vega light just going to be clear about what each of these kind of does and what it solves is sort of a higher-level version of this that really is better for day-to-day usage but still very highly leverageable and in particular where where it differs from from Vega it tends to be in a direction that gives you a lot of flexibility for a day-to-day usage and finally here I've been working on a little library called Oz which lets you create Vega and Vega light specifications on the the closure repple or our enclosure application and then send those over a WebSocket to to a browser and for visualization it can also handle hiccup so that you can create sort of notebook or document like dashboards and have have sort of more control over over things that way and finally it's got a reagent component API so that if you want to be doing kind of dynamic front-end work with this stuff you can you can doing that as well so if you're if you're interested please check that out there's also a couple of talks here linked to if you go to the the github page here one of which is from the interactive data lab folks in Seattle who who built fake and big alight and another is from a recent talk that I did a Seattle closure meetup so closing case here conclusion for data science there's really what I what I see here is a unique intersection of strengths right the data-driven functional philosophy concurrency parallelism distributed computing these are all real strengths of the closure ecosystem this growing sort of trend in the direction of machine learning is really promising and the fact that we actually have a robust front-end target I don't know if I mean someone can sync and pick some bones with me on this after the talk but I don't know of another language that's as well positioned to do both kind of the core work of data science and machine learning that can also do interactive data visualizations or web applications on the front end this is really kind of a unique positioning for for closure and closure scripts and finally I think there's a growing community I here feels felt like a number of years ago you know they're folks doing kind of big data processing and such but often folks wouldn't be talking about that as much in terms of data science and now I see more and more that a lot of the the the companies that are using closure are doing are doing data science work with it and so I think that that's um that's that's a really promising trend finally last note here is that pols his open source software we realized early on that the civic tech community was not going to embrace us if they couldn't sort of see transparency in it and actually you don't want code in your government they were don't really know what's going on and what it's doing and so the fact that this is open source I think is um is a real strength that we have but also what that means is you know if you're this is exciting to you this is sort of like piques your interest and you'd like to contribute in some way or another please feel free to to get in touch and that is all I have thank you [Applause]
Up Next

Radical Transparency and Digital Governance: Taiwan Digital Minister on COVID-19 Response
@PdisTwGov
198 views•2020-06-08

IFS Therapy Demonstration: Complete Session with Unburdening
@IFSCA
95.9K views•2021-01-13

FastAPI vs Flask vs Django: Choosing the Right Python Web Framework
@TechWithTim
302.5K views•2024-05-26

Game of Thrones Opening Credits: A Cinematic Analysis
@gameofthrones
46.3M views•2011-04-18
Related Study Plans & Knowledge Roadmaps
Structured learning paths in General & Interdisciplinary Studies




![[s2 | 2026] Парадигмы программирования, Г. А. Корнеев, лекция 8](https://i.ytimg.com/vi/upWvBnzsIEo/maxresdefault.jpg)

































