Transcript#
This transcript was generated automatically and may contain errors.
Hey there, welcome to the Posit Data Science Hangout. I'm Libby Herron, and this is a recording of our weekly community call that happens every Thursday at 12pm US Eastern Time. If you are not joining us live, you miss out on the amazing chat that's going on. So find the link in the description where you can add our call to your calendar and come hang out with the most supportive, friendly, and funny data community you'll ever experience.
Can't wait to see you there. And speaking of people who make impact, I would really like to introduce you to our featured leader today, Vytas Vaiciulis. He is Lead Data Scientist at CSO, which is Central Statistics Office Ireland. Vytas, thank you so much for joining us today. And I would love it if you could introduce yourself, tell us a little bit about what you do at CSO, and then also something that you like to do for fun.
Thank you very much. So very happy to be here. So I'm originally from Lithuania, but I've been living in Ireland for more than 30 years. And what do you do at CSO? So I lead a section on analytics and AI innovations, which I'm very proud to say, when I began working in that section, it was team of one, and now it's team of 13. So I build it from the ground up, and the entire idea of the section was to modernize our statistics, which I have done, and I'm very proud of it. And that's where my interaction with POSIT comes in, and RStudio. So that's kind of like my professional work. For fun, I have access to the Atlantic Ocean. So I go for swims in the Atlantic Oceans, and in my spare time, I lift heavy objects, and I put them back down. That keeps me sane, because it frees up the mind from my day-to-day thinking, because in data science, there's a lot of thinking, so sometimes I want to forget a little bit, and de-stress.
I don't think you're the only lifter here. I used to lift too. You're in good company.
How Vytas got to Ireland and CSO
How did you get to Ireland? How did you get to CSO?
So I suppose, initially, it was my family. So when I was 13 years old, my parents moved to Ireland. And when I was in Lithuania one of the days, they rang me up and said, son, we're moving to Ireland. I was, OK. So that's how I initially moved to Ireland. Long story short, after that, I did all of my education in Ireland. I went to school, which was a very new experience. And then after I finished up my university, I was applying for jobs. And by pure chance, I ended up in CSO, Central Statistics Office. And I've been always working here, because I always wanted to work with data. It's funny how life works out, because one of the days I was in Lithuania, I was walking around, and I saw Lithuanian Central Statistics Office. I was like, I would like to work for this. And chances happen, I end up working for Ireland's statistical office. But I keep close relationships with Lithuanian statistical office as well.
So actually, my education is quite different. My dream was to get into programming. So that's what I actually studied in the university. I studied computers, and my background was programming. And then when I did master's, I did it in data analytics and machine learning. That's how I developed real passion for AI, as I told people way, way back then when AI wasn't popular. So now when actually AI became popular, I must say, I got on that train right away, because that was my passion all along. It just was sleeping there. The moment wasn't quite right.
What CSO does and the data it works with
Just to give a little bit of context, CSO is a central statistics office for the Irish government. And every government in the world, they have the same, they have a central statistics office. So what we do, we do public releases on various things, like how many people live in the country, which you might know is census. And then we do lots of important economic things like CPI, GDP, things along those lines. And if you combine all the releases together, CSO as an organization does around 300 releases a year.
And the way we get our data for those releases historically has been surveys. So somebody knocks on your door and asks you to complete a survey. That's historically how we get the data in. Now, in the recent times, all statistical offices around the world, they see that people no longer like to answer doors. It's a different area that we live in. So statistics are changing a little bit, but I guess alternative is there. So statistical offices are now moving more to administrative data sources, basically data that does not come directly from the people. So we'll work closely with other governmental departments, with LinkedIn, with all the governmental agencies, and we get data from them, which is transactional level data.
We also see some very interesting sections set up. I even see in my organizations, geospatial data. That's very new. So like in some cases, we have some mandatory statistics that we need to produce for sustainable development goals. And one of them is CO2 emissions. Now, question is, how do you measure CO2 emissions? You can't ask a person, what's your CO2 emission in your backyard? Satellite. Satellite sees that. So I suppose satellite will be one of the most advanced data sources, but just to say the variety of data sources that are coming in moved way beyond just knocking on the doors and asking for people's information to all these diverse data sources, other governmental departments, satellite data, and many others.
Programming languages: from Java and SAS to R
No, when I was in school, I developed a passion for Java. I really liked that. But to be honest, it set me up for life because I really liked the object-oriented way of thinking. But then in practice, once I moved into office, Java wasn't really used. Now, in our office, our legacy language for around 40 years, I would say, was SAS. But we're moving away from that. And that's actually the whole purpose of my job is to move to a more modern language, which is R. But because of that, that's why I built up my section, because we can't just say, oh, we want to do R now, which we are doing it. But infrastructure and software needs to be built up. And that's what I did in my office. I built up the capability.
So now we're fully into R. And if history is to teach us anything, the previous language was SAS, now R. So we're here to stay. And just to say the reason, because I also did the research when I actually began working in my organization, I was tasked with the task, go and research what is the best alternative language. And I did the research. And it's actually on my source of research that we decided to go to R. Let's just say the only two contesters that really came into play was R and Python. But R won it overall. But I see Python bidding in the mix. But I say, let's just deal with one language, please. When you deal with official statistics, the methodological consistency is very, very important. Just say, as an example, we have vital statistics that's actually a time series that goes back over 100 years. So you really don't want to mess with the methodology. You want to keep it the same.
Growing the team and showing the value
The pushback was not major, but I guess from the very beginning, the idea was floating there in the ether. I just sort of took the torch and I showed the value. Okay, so if we move here, what is the value? And I think that's what I've been doing ever since, showing the value. So show them the value, why we're doing this.
And once we began moving to it, the reason why my team grew so much, because the management board, the senior management, well, they've seen the value. I never asked for more people. I was always given more people say, you're doing a good thing. We want you basically to scale. Some people say, oh, can you duplicate yourselves? Like, nope. So therefore we'll give you more people. So it was fairly organic, but there was a little pushback every now and then, but my strategy was always show the value, show why we're doing this. And therefore the pushback was minimal because that's the general trend within the statistical offices.
I just sort of took the torch and I showed the value. And I think that's what I've been doing ever since, showing the value. So show them the value, why we're doing this.
Working with geospatial and satellite data
I manage analytics platform that does statistics. Now, the GIS has become so big that it's entire different section that got set up within the methodology division. I say like the area is so big, I don't deal with it, but I would need to deal with it indirectly because I manage the statistical platform. Because it's becoming such impactful thing, entire new section is now made. And a lot of it is too complicated for me. I understand at the high level, but the things that you mentioned, my colleague, Justin, he also mentioned that Copernicus, they get the data, they try to measure it. And there's a lot of international work that I know that my colleague here does here on the satellite data, because there's a lot of things that you can measure from the satellite data that you cannot do otherwise. Now, my interactions, what it would be when he wants to process large amounts of data, and if he crashes his session, because the data moves into terabytes, that's where I come in. But it's getting bigger. It's definitely getting bigger and bigger. And there's a lot of EU projects and UN projects just dedicated to satellite GIS data.
Administrative data and sharing across organizations
So yeah, that's where I suppose it gets a little bit complicated. We are a governmental agency and work with other governmental agencies. It takes a lot of time and relation building to acquire the data if it's within the other governmental agencies, like it's easier. And we always try to make it very clear, look, we are statistical and that will produce statistics for the greater good. We'll make it all public for the greater good. But again, I don't deal within that area, but I know that it takes a fairly long time to acquire. But once you acquire it, you can produce mountain of statistics. And as I said, different data sources come in. So there's a scanner data from the shops, which is very rich data source for things like CPI and GDP and purchase of goods.
Sentiment analysis and free text data
Short answer, no, I have not heard being used. I have used it in college. I remember it was one of the first projects, but no, we do not use that. We are governmental agency, extremely regulated. The statistics that we do, as I said, we do 300 releases. But no, never sentiment analysis because I could see how that would get, that actually goes back like, how do you clean your data? How a lot of ambiguity could be built in there. And like, no, we try to focus as much as we can on pure facts, statistics, like things that are extremely quantitative, 5.2 million people in the country. So we try to do high, high accuracy as much as we can.
Yes, yes, we do. I think this is for job classifications. When a job description is given, like what is your job and you describe it, like I work with computers and servers and yada, yada, yada. What we need to do, we need to take that free text and classify it against NACE 2.2, which is a big classification that has a bucket for each type of job. So if I work with computers and servers and so on, in which bucket would I go? So the only way that we can classify, say, okay, based on this description, you go in this bucket. So yes, we do free text. And I guess for this particular use case, that's where actually, as I said, vanilla machine learning, that's where it actually comes in very handy. Although a lot of people now experimenting and finding that GenAI is actually pretty good at classifying which bucket which classification you go in based on the description.
Open source adoption in government
Again, it's kind of easy for us that Open Source is being adopted globally and I work a lot with UN and I go to Geneva and I know that there is a project at such a high level that really embraces Open Source at a very, very high level, so that helps. So I suppose like you know it pushes from top down and even within my organization Open Source is now being embraced, when I began it wasn't the case.
Because at the beginning when if you want to do Open Source the main question is the support, we cannot deal with something that doesn't have official support. So again I showed the value and that's where I suppose again Posit came in, it's like we want to use Open Source but we want support, because we cannot support it ourselves but then we want flexibility of the Open Source. And that's where like all the R libraries come in and I guess also even at a lower level we'd be using Linux, Linux is Open Source but we want support so therefore we'd use Red Hat because they have support, that's the underlying operating system that runs everything. So to answer your question I show the value again, I show the value that's what we're getting but what helps is the buy-in from top level which is I suppose like at the UN at the lower level these Open Sources being really embraced so it helps a lot although it wasn't there but now it's there.
The SAS-to-R migration at scale
A lot, a lot, a lot so we are now we are currently organization of 1200 people and I would say people that you know program and use previously SAS day-to-day would be, and this is my estimate it'd be like 500. So a lot and I am I'm estimating I say like in total have 1200 people I guess senior management you know they don't really use it but day-to-day people that would be running programs and then we'd have around 200 statisticians, now statisticians this is where your like heavy programming lies. These are the people that write all the statistical releases, they come up with methodology, but after that like nowadays like you know other people as well supporting admin functions they would also be so in total I would say around 500.
Now I have pushed more than half already into the new platform and just to supplement that like for us it's a five year project and we are more than halfway through. This is the biggest project for our organization which I happen to lead and the biggest amount of elephant in the room is the code. The infrastructure is built, expertise is built, we're very big fans of git and RStudio integrates with git very well so that's one of the big deals. Again I show the value and great tool speaks for itself, you know it doesn't need advertisement. Like I would say like you know the git was very very good because once people put their code on git in internal git, now need to mention internal version control.
Using Shiny at CSO
So I suppose just to give, when I built it within my organization I called our platform, so we use everything with the workbench, the package manager and the Shiny. Now the Shiny has become too popular and that's why, that's why my team have grown so much. So we use it a lot, we have a lot of internal apps and yeah we use it a lot a lot. There's a lot of internally facing us because when people develop Shiny apps there's a skill element because I always say like no, before let's say if you're not using Shiny what was the alternative, you become a full dev developer, you learn like html javascript css and nobody really knew that. But now because we have the connect server anyone can spin up a Shiny app and it's very easy and that's where we see like a lot of innovation and then when you combine that with AI well that's another level of capability.
And it's very easy and to be honest it makes it too easy. Too easy. I actually have to say no to people in some cases because we use it a lot and we have many many Shiny apps. It's too easy to make a Shiny app. We actually developed a Shiny dashboard to show how many Shiny apps we have and I think it's very meta, it approached three digits.
It's too easy to make a Shiny app. We actually developed a Shiny dashboard to show how many Shiny apps we have and I think it's very meta, it approached three digits.
AI governance and the challenge of integrating AI
It's extremely cautious and regulated so quality and safety is always first. So I suppose like we are proceeding very cautiously. If I want to do AI it takes a long time to get it approved because there's so many assurances that you need to and rightfully so. I want to implement AI and I need to jump through many hoops but to be honest I do it with pleasure because I well not with pleasure but I understand the reason behind like you know everybody needs to be assured. And we have AI governance board within my organization that involves everybody from all over I suppose like aspects of the organization, you know HR, IT infrastructure, everyone. So everybody wants to do AI there's dozens and dozens of use cases but everybody's being very cautious and at the center of it being data.
So it's an ongoing thing but even at the high level at the senior management level like the language is changing. I actually began, I developed a use case for AI which was a SAS star in 2023 which created us immense value and back then they're like what are you using an LLM, what's an LLM, nobody knew back then it just came out. But now everybody knows it, again it comes back to value, can you show the value, why this is useful, if you can show the value useful well that helps and that's my been I suppose that's the way I operate. I show the value explicitly clearly and that gets the buy-in.
Maintaining methodological consistency during migration
That's why the project is five years, that's why the project lasts five years because as I said before all of our statistical releases are underpinned by code. Now all of this code needs to be translated, when we translate the code we need to maintain methodological consistency. Now we have multiple ways to address this. I would say one of the most useful ones we do parallel runs so we have a SAS and we have R writing our code, we use historical data, run it at the same time, produces the same results. Now every statistician is an expert within their domain and it's their responsibility to do that but we also do code reviews on top of that where methodology quality it would come in. But that is the reason why it's taking us five years. Now we made a very good progress and we'll keep on moving on. So it's a parallel runs, code reviews by other than the statistician that owns the code because you just need somebody from outside to have a look, have a sanity check.
Career advice
I would say like as I grew I changed, I did not remain the same as I when I began my education. And I say it might seem a bit stereotypical, don't focus on the tool. Like you know I say like the thing that helps me the most it's kind of like logical reasoning, breaking problem down. So whatever the tool is it's kind of like you know learn how to solve problems, do not shy away from a challenge. Because even within my organization when I began working like you know tools change and technology changes but within the area that I worked it's always like you know fixing a problem. So logical thinking, breaking down the problem, I think like that's the thing that helps me the most.
And in addition to that the thing that I neglected the most when I was in college and I thought that it's not important at all it's like I don't even want to do this module which is what I'm doing now, it's communication. Being able to talk to people and communicate well because you could be the smartest person in a room but like if you cannot communicate that, if you cannot explain it, like in a way you don't scale. So also explaining complex matters to people that's serving me a lot. So breaking down problems irrelevant of the tool, these days AI is a big thing, and then communication. A lot of it's like ah no communication come on man like you need to have heavy skills, yes you need to have heavy skills and understanding of technology but technology changes, if you look back at the history communication I don't think it's gonna go away that time soon. And that helped me the most, you know that's why I went to UN meetings and Eurostat and all these places because I said okay I think this guy understands what he's talking about so let me speak.
Being able to talk to people and communicate well because you could be the smartest person in a room but like if you cannot communicate that, if you cannot explain it, like in a way you don't scale.