Transcript#

This transcript was generated automatically and may contain errors.

Hello, everybody. Welcome to the Data Science Lab. My name is Libby Herron. I'm a Data Community Manager here at Posit. I'm joined by Isabella Velazquez, my co-host today, and your maven on Discord. Isabella, would you like to say hello? Hi, everyone. Thanks so much for joining us.

We're so glad you're here. And we are joined by our lab manager for today, Barret Schloerke. Barret, would you like to introduce yourself? Hello, my name is Barret Schloerke. I am on the Shiny team at Posit. And yeah, just discovered Worktrees, you know, a few months ago and haven't looked back. So I want to share with everyone all that I know.

Yay, that's what we're all about. Okay, if you've never been to the Data Science Lab, we get together every Tuesday, same time, same place, most Tuesdays. Sometimes we take breaks and sometimes there's holidays. But we are here to do exactly what Barret just said, to share knowledge, to talk about technical content.

And this is a place where we do not have a bunch of slides and just talking and lecturing and presentations. This is more about screen sharing and hopefully doing a little bit of like exploring and paracoding with your friends on the internet, your data friends on the internet.

Introducing git worktrees

Today, we are going to be talking about Git Worktrees, which is a concept that did not come into my brain until, same as Barret, a few months ago. Isabella and I were working on the Data Science Lab episode that we did for Git and GitHub integrations in Positron, which was very, very fun and super, super great. I think a lot of people had fun with that one. So look for that one on YouTube. But inside of our exploration in Positron, I remember I kept seeing work trees over and over again. And I'm like, what the heck is a work tree? And then I noticed more people talking about it. You know, when you hear a word or you see a word, you learn a word, and then you see it everywhere.

That started happening with work trees. And then a lot of people in the lab started asking about them. What is a work tree? So Barret is here to talk about them. Barret, I am going to give a brief explanation, and I'm going to let you take it away with some demonstrations of what I am talking about.

If you have never explored work trees before, work trees allow you to basically check out two different Git branches at the same time locally. So usually with Git, you check out a branch, and you have a view of your project, right? You have a view of all of your folders inside of your project. And you can make changes, commit them. But if you haven't really like finished what you're doing, you sort of have to like stash your changes or commit it temporarily. And then you can go check out a different branch and do other things. This allows you to check out two different branches at the same time locally, which you can switch between, but they are visible to you in a folder structure. So each version of your project basically is in a different folder.

It allows you to check out more than one branch at the same time. They have to be different branches. They of course cannot be two versions of the same branch that would create a local conflict that would not work, and then the universe would implode. So that doesn't work. They have to be different branches.

No, that's good job, Libby. The gist? That's the gist. Yeah. Okay.

Exactly that. So yeah, I will reiterate, WorkTrees are, it's as close to a full clone to a new repo. So that, except the only thing with WorkTrees is that they try to make it a, it's not done ad hoc. You're not just willy-nilly checking out, cloning your repo into different folders. It's a organized way to do so. And there's ways to do it on command and also to clean it up. And cleaning it up is important because you don't want just clones existing forever and ever and ever.

I think branches feel more like diffs. So it's just little alterations where WorkTrees on your hard drive, it's physically two full copies.

where WorkTrees on your hard drive, it's physically two full copies.

Yes. So instead of in the mental model of branches, where a branch is only storing what has changed, right? Because it's like, I'm only going to look at what's changed. Everything else is the same. We're going to leave it. We're only going to store the diffs, which is what you are seeing when you're looking, you're looking at the differences. But with WorkTrees, it is like, we're going to bake you a whole nother cake. You have a whole extra thing that you are storing.

I think that's a really, really important thing to remember. And I saw on, I want to say BlueSky this morning or last night, Carlos, who is also at Posit talking about how he had filled up an entire two terabyte hard drive or something with WorkTrees. I think they can add up very quickly if you're doing a lot of things at one time.

When worktrees are useful

So Barrett, when is understanding and using WorkTrees important and valuable? Because I've never used it before or needed it before. Obviously, I haven't hit that use case. But the way that a lot of people are working every day, especially using AI agents, is changing what people need.

Yes. So short answer is, whenever you would say, you know, get checkout, make new branch, whatever incantation you wanted to do, I've replaced that completely with make new WorkTree.

And this has removed my need to know about stashing or like committing temporarily. I don't need to do all of these little workarounds to get my job done. And I think it's really beneficial because like, you could be two days into a feature and then your boss or, you know, CI comes back and says, you need to fix this now. And like, it's so disrupting to try to figure out like, how do I pause this to go over here? And instead with WorkTrees, it's like, all right, from main branch, make new WorkTree. And it's in a different folder.

So you can have, you know, kind of how when you have your RStudio IDE or Positron, you have like your unsaved files, and they just magically reopen when you open up your project again. It's like having that with WorkTrees, you could have your unsaved changes. And it's separate, it's different. And it works out well that way. And so anytime that I would say git, you would have a git checkout to start a new feature. Instead, we should create a new WorkTree. And then we can do our work in parallel that way.

Okay. And I think it's important to say that this type of working will not be for everybody. There will be a lot of teams who are just like, we're going to keep doing exactly what we're doing. We don't have a need for WorkTrees. This is a place where we can just be in receive mode and learn about this stuff, even if we don't need to directly apply it right now.

Introducing Conductor

And Barrett is also going to show us the sort of next Pokemon evolution of that, which is orchestrated WorkTrees, which is a WorkTree first way of working. He uses a tool called Conductor. And I will let him explain a little bit more about what Conductor is, because I think that it's going to be a little bit of another level of brain expansion.

Exactly that. So yeah, I will reiterate, WorkTrees are, it's as close to a full clone to a new repo.

Conductor.build here. It is a Mac only application. There is a there's another application called Superset. It achieves the same goals as Conductor. I just have enjoyed Conductor because it's a little more UI heavy instead of where Superset just kind of feels like terminal coordinator. I like a little bit more UI in my applications. But Conductor, it allows you to run a team of coding agents for your Mac. And that sounds very intimidating. But the way I interpret it is that with the click of a single plus button, I can make a brand new work tree with his full file copy and a full everything.

This sounds like a nice way to teach actually from a real world project because you can be like, well, I don't want to mess my project up by showing all of my learners this stuff. Let me make a demo work tree real quick. I can destroy everything. And it won't touch my actual project. Then I just delete the work tree.

Yes. Because oh, one thing that we did not mention was that you merge just the same way as you would a branch with a PR. So that's how you can sort of resolve your changes that you've made in a work tree into your main project.

Yes, exactly. And because it's using the familiar merging, this is why I've switched all of my create branch workflow with no create work tree.

So now I just hit my plus button, and then I'll create a PR in the top right corner for me. And what also might help us understand is Lauren's question in the discord, which is, where does it make this copy? Where does it live? Is it in the same working directory?

Absolutely. It's great. Very important. In general, I went through my settings earlier, just to make sure I wouldn't show you my API keys just by accident.

So the root path ends up being this is the repo. And then the workspaces go inside my like a conductor parent folder workspaces. And then every single of my local projects on the left will have their own folder here. And then inside that will be like, kind of like the branch name, but it'll be the work branch or the work tree name. So if I just name your work trees after your branches, but you can call them different things. I've seen people call them the name of the agent that's working on it. I've seen them call them just like fun things.

Agentic workflows and parallel worktrees

Speaking of hopefully, why work trees are useful for me, as well as there will be times where I may send an agent on a task. And it may work for it. It's not it may work. I routinely have 15 to 45 minute run times that my agent is working. And I don't know when it's going to finish. I don't know, you know, when it's gonna be ready for me to help it. And so it's implementing a long time. I have some co workers where they let it kind of run for like four or five hours, and then they check on in on it in the morning, and they check on it at the end of the day.

And as a background task. Wow, so the work tree is working on in parallel, just the background, and you can do other stuff on other work trees in the meantime. Yes. Yeah. So since I'm twiddling my thumbs for 15, you know, 30 minutes, like, yeah, I not start the same process. And now I can, you know, if I feel like I'm spinning plates, all at the same time, plates that won't crash into each other.

This is another thing that somebody brought up on blue sky when we were talking about it, was that it creates this sandboxed environment where your agents cannot step on each other's toes. They cannot like bump into each other by and cause a cause conflicts because they're in totally different spaces. And likewise, it can't bump into you or step on your toes.

So, RoboRev is a, as the name kind of hints at, it's a robotic code reviewer. And what its job is, it's trying to do, is that whenever you make a commit, you can do this locally or on GitHub. But whenever you make a commit, RoboRev will insert itself into the process and say, let me code review this right now. Not after you've made 10 commits. Because if I'm, for instance, today I'm using Claude. Claude will perform 10 tasks, and so it will commit along the way as it's executing its work. And then when it's like, along the way, RoboRev will inject itself saying, hey, I'm a code reviewer, let me work in here, please. And it'll be like, yes, I agree with the decisions you've made so far, you know, continue. Or it will say, oh, linting errors, please fix this. Or I found a security flaw, fix it now, not at stage 10, but at stage 2. So, it tries to catch things and steer things early.

WorkTrees without agents, and use cases

hey, so WorkTrees are specifically useful for when you're using an agent workflow, or can they be useful without agents? They can be useful without agents. I think they are more useful by far if you are working in an agentic workflow. Barrett, does that seem right?

Absolutely. Yeah. Okay. When there's this waiting time that you can't control, it's much more useful with agentic workflows. But I would personally just switch all my branch workflows to be WorkTree workflows instead. I think it also benefits you more when you have a robust and large code base. If you have a very, very tiny code base and a very, very tiny project with very, very few possibilities of this multiverse situation, it's probably a little bit less useful, and it might be overkill, right?

I have also seen people use it for testing. They might sort of A-B test something where they're like, I have two implementations of something on my repo. They're in two different WorkTrees, and I'm going to test each of them for efficiency and for all kinds of different things to see which version I like better and which version works better, and then keep the one that works better. And running tests can take a long time, so you could run them at the same time on different branches.

Conductor walkthrough

Okay. Let's get back to Conductor and a little bit of a practical explanation of what we are seeing on the screen. So, my Git login is off. I don't know why. It is working in the terminal. So, I can just make a new PR, like normal, rather than having Claude try to do it for me, and it would be there. So, I didn't have the agent make it for me, but in the past, I think I've made like under five commit messages or PRs manually in the past six months. I've switched over so hard to having agentic workflows, just because it works for my use case.

Conductor will tell you that, hey, a draft PR has been open. Here's the PR number. It's ready for review, and we can get that merged.

it is as easy just to say make me a PR. And it'll tell the agent to write up instructions, has a default instructions that it gives.

I can say merge. So, I'm going to do that, and it will actually come through, merge it, I hope, unless, yeah, and now it archived my workspace. The PR was merged, and my workspace was archived, so it auto cleaned up everything for me, given that PR was merged, and now I just can come back and hit plus and move on, and we'll actually see that in the read me at the bottom, authors, all right, it is there, because it bases everything off of origin main to begin with.

So, yeah, it's just, feels very similar to, like, if I didn't tell you this was work trees under the hood, it was just branches, you would feel like you're none the wiser about what's actually happening.

if I didn't tell you this was work trees under the hood, it was just branches, you would feel like you're none the wiser about what's actually happening.

Right, because instead of you actually looking at the folder structure separately, when you're using Conductor, it is just like you're in a, you're just checking out different branches.

Data science use cases

One is a follow-up from Matan that says, is this workflow, do you think mostly for software development, or could you think of use cases for data science or analysis or something similar?

So, my wife is a data scientist. She's at Sleeper, kind of like FanDuel, and they've been doing some database migrations, and that's been taking three weeks, you know, they're doing multiple PRs of files to be migrated, and definitely should have listened to Michael Chow's DBT talk recently. That would have been very useful, and so this is, you know, a very long process that's happening. You could do this in serial branches or serial, you know, PRs, but also she needs to do some analysis at the same time, and those are going to be three or four day projects that are needing to be happening at the same time. So, she can come over here and say, okay, I'm going to do this, I'm going to do this, I'm going to do this, and it's going to be happening at the same time.

So let's, you know, migrate these files, blah, blah, blah, and then come here and say, do analysis on this data, and now they're separated into different areas, and she doesn't need a stash, she doesn't need to do anything, and when one of the migrations need a hot fix right now, she can go do that immediately. Without having to temporary commit her changes, or stash her stage changes, or any of the things, she can just move her focus and then move it back. Move to the stove on the right while the steak is cooking, work on the risotto, and then come back.

IDE integration and worktree-first thinking

Absolutely. Okay. What I really want is this left sidebar to be an extra sidebar in all of the IDEs. And this is another one on the left. This is, it's a hard one, because Positron, and Zed, and VS Code, they look at this folder, and this folder is narrowly scoped to only this folder. It doesn't work for the case of working trees, where I'm saying, I want to belong to, in this case, the Shiny React repo, and within the Shiny React repo, I want 10 different parallel work trees, and I want to switch between them like conversations. And within that conversation, I want a whole VS Code, I want a whole Positron, I want a whole Zed. That's what I want.

As Daniel said in the early message, VS Code and friends are not work tree first. They are a follow-up to making a work tree.

Yes, exactly. I think it's a different sort of way of working, way of thinking, way of approaching everything.

Okay. We had a question from Nathan that said, I'm wondering if work trees are a way to get around rendering multiple Quarto reports. I tried rendering them in parallel last week, and they all crashed because they tried to write temp files to the same folder.

It depends on your Quarto documents. If you're checking, if you're able to check in the rendered results, then sure, you can do that. But because if that's the case, I would argue that your documentation is a one-way trip and instead you should have CI render your documentation for you. And so your machine is free from having to do that ever and have CI commit it back to the repo when it's done.

CI, like continuous integration? So you might have a CI pipeline. I always try to define acronyms for people because often in data science and in software engineering we just use acronyms like they're the word but we don't explain what they are.

Continuous integration, continuous deployment, I think. Yes, exactly. So as let's say it was this one and inside the commits there would be a workflow in GitHub that would come in and say, set up job, check out the repo, run Quarto document, commit results and then it would be done. And then it would commit back to your PR or back to the main branch.

Using worktrees without Conductor

Another question from Anna. Anna, you're asking great questions today, thank you. It says, do you need something like Conductor to work with WorkTrees or can you just use it in your terminal, especially for non-agentic workflows? You do not need Conductor. In fact, you can do it in the UI in Positron which I think somebody probably replied and answered in the discord.

Yeah, if you are in Positron, you can go to your little three dots next to your branch and you can click that and go down to WorkTrees to create a WorkTree from the branch that you are on and you can name it. You can also just do command shift P, open up your command palette, type in WorkTrees and you will get your little three options there whether that's create WorkTree, delete WorkTree or migrate, there you go. You can open one. So you can do everything that you need to do in a WorkTree without a Conductor.

But I get the feeling that if you have more than one WorkTree on the go, right? Like you have the need to create many different stoves in your kitchen, it's probably easier to have somebody that you can put in charge of those stoves, right? Yes. And that would be Conductor.

Yeah, if you are terminal friendly and you want to use the commands, there are full commands within the terminal for setting up WorkTrees.

Memory management and archiving

Is there a good way to keep track of how much memory you are filling up with all of these things you're creating? And like, when do you know it's time to delete them? Do you, do they, you have that set so that they archive upon merge, right? So like you do a PR, you do a pull request, you merge into your main branch, and then it archives them. Does that mean that it is no longer taking up space, that it's minimizing them in some way? Or do you still need to go manually delete them?

It, I think, will delete it after 30 or 90 days if it's just archived, kind of like Gmail does with your trash. It doesn't get rid of it right away. But you can, in the settings, set that delete branch on archive. And if you delete branch on archive and you enable this one, it will make it so that when you archive, it is gone, gone. It's not archive, it's delete. Basically archive equals equals delete at that point.

Do you ever go back to ones that you've archived and reopen them? I have. And that's why I do not have that one checked anymore.

So I think we added this one. And if I move Zoom, this was the one that I added my name in the readme. And we can see that the PR is merged. But if I come back, we have the full chat for telling me that I couldn't do my login anymore. And it's all there, just as if we never left. And so that's why I like that archive. And then, you know, my attention spans, if I'm not back to it in 30 days, I don't need it, so.

Conductor's city naming game

There's something with conductor where they have a mini game of, if you open up enough work trees, they didn't have it right away. But if you open up enough work trees that it will, yeah, 74 of 307 cities visited. And not all cities have the same proportion, but it's like, oh, I was just in Memphis. Okay, well, somewhere over here, I visited Memphis. Like, ta-da. So it's picking city names to name your work trees.

It needs that temp name to make the work tree. And then once work has been started, it makes a new name that points back to that temporary folder so that it doesn't have two checkouts.

Yeah, so that's what it does. Because like a git branch checkout, you need a name to check out into first. And so since it doesn't know what you want to do, it just kind of makes a fun city name.

Forking worktrees and the concept of main

So yes, there is a trick here of fork to new workspace. So I have my WorkTreeFunAnalogies task here, and I can fork to workspace. It makes a new one, and now I'm technically in Wet'suwet'en, and this one is just actually WorkTreeFunAnalogies, and it says, please continue from this one with a summary of where we were, and so it essentially has forked my WorkTree and Claude session over.

And now we can continue here, but you get in this weird spot of likeβ€” This is WorkTree Inception.

There's origin main, but don't worry about local main. I think that there needs to be a description of that. I know we only have five minutes left, but when you are working on your computer, you are working on local versions of what exists remotely, for example, on GitHub. It doesn't have to be GitHub, it could be GitLab, it could be something else. Did you are working on something that is a local version of that?

That's why if you, like, let's say you've made a commit, or let's say you've merged something, you've merged a branch into main on the GitHub UI, you'd still don't have a version of that main on your local computer. You have to pull that down and sync your changes so that your version of main now looks like what the one at the origin looks like, right?

So really, main is just another WorkTree, it's just like a version of your origin. So what Barrett is saying is, we don't ever go check out main, it's just a view of your origin. And so all of these are just views of the origin, but they're actual copies, because they're clones.

Before WorkTrees, I actually saw Joe Chang have six or eight different copies of R Shiny on his computer. Because he would just check out, he's like, oh, I don't wanna like deal with this right now. So he would just do a git clone into a new folder, and like, he would just manage it manually every once in a while. I don't know how he kept it straight in his head.

But yeah, before WorkTrees, it was just the Wild West.

Wrapping up

So Tom asks, so main is only relevant when you merge, and it's a moving target otherwise? Yes. I think, yeah, I think that's a good description.

I hope everybody is a little bit closer to understanding what a WorkTree is, the utility of WorkTrees, the use cases for WorkTrees, and maybe a little bit of how you might need them or utilize them. And I really recommend, and I was talking with Jenny Bryant about this too, the best way to play with these is just to go open something that doesn't matter and play with it. Create a WorkTree, do something with it, merge it, nonsensical changes, just open up text files, it doesn't matter. See how they work, look at the folder structure.

So my origin one here is inside my, you know, RStudio Shiny React, Shiny React NoSync. Like that's my original. And then all of my WorkTrees are inside Barrett Conductor WorkSpaces Shiny React. So they're all over here, full copies.

And then Nathan had asked, are we operating in a world where you don't push directly to main? You just make a new branch. You still do PR to main. You are PRing to main origin, which is up remotely. So that's a fair one. If you were to do work in main and then boss comes in and says, you need to fix this right now. You have to clean up that working area, that stove, that kitchen stove. You have to clean it up before making your fix. If you were on a working branch or a WorkTree, you could do that just freely. You would just make a new one, make the fix and merge back to main. So personally, make small PRs, make many, many, many, many PRs and merge often and quickly. And WorkTrees allow for that.

Amazing. Barrett, my brain's exploding in the best possible way. I hope all of you have learned so much and are thinking about ways that you can take this forward. Everybody, big round of applause for Barrett.