Identifying and tracking toil using SRE principles

One of the key measures that Google site reliability engineers (SREs) use to verify our effectiveness is how we spend our time day-to-day. We want ample time available for long-term engineering project work, but we’re also responsible for the continued operation of Google’s services, which sometimes requires doing some manual work. We aim for less than half of our time to be spent on what we call “toil.” So what is toil, and how do we stop it from interfering with our engineering velocity? We’ll look at these questions in this post.First, let’s define toil, from chapter 5 of the Site Reliability Engineering book:“Toil is the kind of work that tends to be manual, repetitive, automatable, tactical, devoid of enduring value, and that scales linearly as a service grows.”Some examples of toil may include:Handling quota requestsApplying database schema changesReviewing non-critical monitoring alertsCopying and pasting commands from a playbookA common thread in all of these examples is that they do not require an engineer’s human judgment. The work is easy but it’s not very rewarding, and it interrupts us from making progress on engineering work to scale services and launch features.Here’s how to take your team through the process of identifying, measuring, and eliminating toil.Identifying toilThe hardest part of tackling toil is identifying it. If you aren’t explicitly tracking it, there’s probably a lot of work happening on your team that you aren’t aware of. Toil often comes as a request texted to you or email sent to an individual who dutifully completes the work without anyone else noticing. We heard a great example of this from CRE Jamie Wilkinson in Sydney, Australia, who shared this story of his experience as an SRE on a team managing one of Google’s datastore services.Jamie’s SRE team was split between Sydney and Mountain View, CA, and there was a big disconnect between the achievements of the two sites. Sydney was frustrated that the project work they relied upon—and the Mountain View team committed to—never got done. One of the engineers from Sydney visited the team in Mountain View, and discovered they were being interrupted frequently throughout the day, handling walk-ups and IMs from the Mountain View-based developers. Despite regular meetings to discuss on-call incidents and project work, and complaints that the Mountain View side felt overworked, the Sydney team couldn’t help because they didn’t know the extent of these requests. So the team decided to require all the requests to be submitted as bugs. The Mountain View team had been trained to leap in and help with every customer’s emergency, so it took three months just to make the cultural change. Once that happened, they could establish a rotation of people across both sites to distribute load, see stats on how much work there was and how long it took, and identify repetitive issues that needed fixing.“The one takeaway from this was that when you start measuring the right thing, you can show people what is happening, and then they agree with you,” Jamie said. “Showing everyone on the team the incoming vs. outgoing ticket rates was a watershed moment.”When tracking your work this way, it helps to gather some lightweight metadata in a tracking system of your choice, such as:What type of work was it (quota changes, push release to production, ACL update, etc.)?What was the degree of difficulty: Easy (<1 hour); Medium (hours); Hard (days) (based on human hands-on time, not elapsed time)?Who did the work?This initial data lets you measure the impact of your toil. Remember, however, that the emphasis is on lightweight in this step. Extreme precision has little value here; it actually places more burden on your team if they need to capture many details, and makes them feel micromanaged.Another way to successfully identify toil is to survey your team. Another Google CRE, Vivek Rau, would regularly survey Google’s entire SRE organization. Because the size and shape of toil varied between different SRE teams, at a company-wide level ticket metrics were harder to analyze. He surveyed SREs every three months to identify common issues across Google that were eating away at our time for project work. Try this sample toil survey to start:Averaging over the past four weeks, approximately what fraction of your time did you spend on toil?  Scale 0-100%How happy are you with the quantity of time you spend on toil? Not happy / OK / No problem at allWhat are your top three sources of toil?On-call Response / Interrupts / Pushes / Capacity / Other / etc.Do you have a long-term engineering project in your quarterly objectives?Yes / NoIf so, averaging over the past four weeks, approximately what fraction of your time did you spend on your engineering project? (estimate)Scale 0-100%In your team, is there toil you can automate away but you don’t do so, because that very toil takes time away from long-term engineering work? If so, please describe below.Open responseMeasuring toilOnce you’ve identified the work being done, how do you determine if it’s too much? It’s pretty simple: Regularly (we find monthly or quarterly to be a good interval), compute an estimate of how much time is being spent on various types of work. Look for patterns or trends in your tickets, surveys, and on-call incident response, and prioritize based on the aggregate human time spent. Within Google SRE, we aim to keep toil below 50% of each SRE’s time, to preserve the other 50% for engineering project work. If the estimates show that we have exceeded the 50% toil threshold, we plan work explicitly with the goal of reducing that number and getting the work balance back into a healthy state. Eliminating toilNow that you’ve identified and measured your toil, it’s time to minimize it. As we’ve hinted at already, the solution here is typically to automate the work. This is not always straightforward, however, and the aim shouldn’t be to eliminate all toil.Automating tasks that you rarely do (for example, deploying your service at a new location) can be tricky, because the procedure you used or assumptions you made while automating may change by the time you do that same task again. If a large amount of your time is spent on this kind of toil, consider how you might change the underlying architecture to smooth this variability. Do you use an infrastructure as code (IaC) solution for managing your systems? Can the procedure be executed multiple times without negative side effects? Is there a test to verify the procedure?Treat your automation like any other production system. If you have an SLO practice, use some of your error budget to automate away toil. Complete postmortems when your automation fails, and fix it as you would any user-facing system. You want your automation available to you in any situation, including production incidents, to free humans to do the work they’re good at.If you’ve gotten your users familiar with opening tickets to request help, use your ticketing system as the API for automation, making the work fully self-service.Also, because toil isn’t just technical, but also cultural, make sure the only people doing toil work are the people explicitly assigned to it. This might be your oncaller, or a rotation of engineers scheduled to deal with “tickets” or “interrupts.” This preserves the rest of the team’s time to work on projects and reinforces a culture of surfacing and accounting for toil.A note on complexity vs. toilSometimes we see engineers and leadership mistaking technical or organizational complexity as toil. The effects on humans are similar, but the work fails to meet the definition at the start of this post. Where toil is work that is basically of no enduring value, complexity often makes valuable work feel onerous. Google SRE Laura Beegle has been investigating this within Google, and suggests a different approach to addressing complexity: While there’s intense satisfaction in designing a simple, robust system, it inevitably becomes somewhat more complex, simply by existing in a distributed environment, used by a diverse range of users, or growing to serve more functionality over time. We want our systems to evolve over time, while also reducing what we call “experienced complexity”—the negative feelings based on mismatched expectations about how long or difficult a task is to complete. Quantifying the subjective experience of your systems is known by another name: user experience. The users in this case are SREs. The observable outcome of well-managed system complexity is a better user experience.Addressing the user experience of supporting your systems is engineering work of enduring value, and therefore not the same as toil. If you find that complexity is threatening your system’s reliability, take action. By following a blameless postmortem process, or surveying your team, you can identify situations where complexity resulted in unexpected results or a longer-than-expected recovery time.Some manual care and feeding of the systems we build is inevitably required, but the number of humans needed shouldn’t grow linearly with the number of VMs, users, or requests. As engineers, we know the power of using computers to complete routine tasks, but we often find ourselves doing that work by hand anyway. By identifying, measuring, and reducing toil, we can reduce operating costs and ensure time to focus on the difficult and interesting projects instead.For more about SRE, learn about the fundamentals or explore the full SRE book.
Quelle: Google Cloud Platform

Introduction to Customer Empathy Workshops

Product feedback from users goes a long way. It’s why Red Hat’s OpenShift Web Console UI is as awesome as it is today. Features like Dashboards and Topology were added because of user feedback—and that’s how we plan on enhancing the console even further. One thing’s for sure: The path to a better console experience relies on continued customer engagement.
Thus, Red Hat has launched a series of workshops specifically geared towards engaging and empathizing with OpenShift customers in order to better understand their needs. We’ve dubbed them customer empathy workshops.
What they are
Our customer empathy workshops enable customers to collaborate with OpenShift user experience, development, and product management to directly influence the future of the OpenShift console. Each workshop has a special topic to focus on so that the group can really hone in on the challenges they face in specific areas. Customers are introduced to our design thinking process as we dive into real product development challenges, starting with problem discovery and following with solution ideation.
The design thinking process

Hands-on activities give our customers the unique opportunity to connect with the OpenShift team, share their pain points, and collaborate with other community members throughout the session. This kind of collaboration makes the product what it is today, so we want to continue engaging with users as much as possible.
Value for our customers 
These workshops certainly help the product evolve, but they also give customers an opportunity to discover, impact, and connect.
Discover: Our customers will get the opportunity to learn how product decisions are made from the small fixes to the larger feature additions. They can also share their pain points, what they struggle with, and where they need help—as well as learn how other companies have overcome similar obstacles.
 
Impact: Customers can lead the conversation around the OpenShift user experience, engage with other OpenShift users, and collaborate through knowledge sharing and group solution ideation.
 
Connect: Discussing the OpenShift console with users brings together folks from different countries, industries, and technical backgrounds. We hope that our participants walk away with new connections and feel even more connected to the OpenShift community.
Our opportunities
While customers are discovering, impacting, and connecting, we’re gaining valuable insight from all the feedback. Specifically, we have the opportunity to listen, prioritize, and design.
Listen: We want to learn more about how our customers use OpenShift: What their environment looks like, how many people are on their team, what their biggest pain points are, and more.
 
Prioritize: Through hands-on activities, we hope to better understand the problem statements that arise throughout our workshop. The more we learn about customer pain points and what ideal solutions might look like, the better we can design a powerful experience.
 
Design: At the end of the day, we want to take all customer  feedback and implement features to make the OpenShift experience better. So after each workshop, we’ll analyze the data, explore the proposed solutions, and design a fix or new feature to address it.
Stay tuned
This series has been an exciting addition to our engagement efforts, and we’ve heard some great feedback from those who have already participated. Following each workshop, we’ll share a summary of the workshop and preliminary results. So keep an eye out for upcoming customer empathy workshops and content. We look forward to sharing the results with you in an upcoming blog article!
The post Introduction to Customer Empathy Workshops appeared first on Red Hat OpenShift Blog.
Quelle: OpenShift

How Omnitracs Transformed to a DevOps Culture with OpenShift

Omnitracs has taken an interesting road to get to its current position as a leader in fleet management software for logistics and transportation companies. Their SaaS-based offering allows companies to track, monitor, and bring into compliance all of their trucks and shipping vehicles around the globe from one system. But just because Omnitracs users were taking advantage of cloud-based software as a service models of consumption doesn’t mean Omnitracs developers were fully utilizing the cloud and the agile methodology it enables.
That’s only been the case for the past year, in fact, since Omnitracs began adopting Red Hat OpenShift. Andrew Harrison, lead IT DevOps Engineer and lead of the Agents of Change team at Omnitracs, was tasked with building the company a road to the future of software development, and the pavement on this road was built with OpenShift.
Since 2014, Omnitracs has been growing rapidly, launching over 30 new products, and merging in the assets from a number of acquired companies. To keep up with all of this growth, the developers in the company had to transform their way of doing things, top to bottom.
Thus, a year ago, Harrison was placed in charge of affecting change throughout Omnitracs’ IT organization. That means introducing devops, automation, agile methodologies, and continuous integration and deployments. That’s a tall order for a single team to spread such changes through an entire enterprise.
And yet, a year later, Harrison said he’s successfully transitioned the company away from a “waterfall” style of development and deployment towards a devops and agile based approach, thanks to the help of the Red Hat OpenShift platform. Instead of code pouring in as it was completed, like a waterfall, developers were able to iterate over time in smaller chunks. While the move began with OpenShift 3.11, the company was also one of the first to roll OpenShift 4.1 out to production systems.
Harrison said the benefits of moving to OpenShift 3.11 were immediately noticeable. “With OpenShift 3.11, we were able to get immediate cost reduction because we moved everything out to the public cloud [eliminating on prem costs]. We got rid of on-premise costs immediately. We wrote custom Ansible playbooks to make sure everything was infrastructure as code and was always repeatable. We reduced our environment deployment times from over a week to less than 2 hours,” said Harrison.
Those technical wins were paralleled by cultural wins as well, he added. “We had a total transformation in the IT department, especially in our organization. Our team, The Agents of Change, evolved from the traditional waterfall style to a real agile methodology which greatly reduced all of our release times.”
Soon after moving onto OpenShift, the team at Omnitracs embraced OpenShift 4.1 and the Operators model for services availability.Harrison noted that the Operators model enabled faster integration of services essential to building out a production-grade cluster. The team now uses Splunk for logging, and Ansible for automation.
Software is not the only thing changing inside the Omnitracs OpenShift clusters, however. The company is also using this Kubernetes-based cloud software to improve its internal IT teams. Said Harrison: “We took the SRE [Site Reliability Engineering] model and turned it on its ear a little. We’re spinning up 35 new scrum teams in the coming years, and we’re not going to find SREs for each of those teams; it’s not going to work. So we took a different approach: we’re using a virtual ops approach, where we’re embedding with these teams now and teaching them the ops part of devops. They are going to be their own SREs. They will have the ability and proper permissions in OpenShift to spin up their own namespaces, start projects, give people access to them, and all of that. We’re really giving them the ability to provide care and feeding for their own areas, but at the same time, giving them the ownership of it. So now when they are building their applications, they are actively thinking about the operationalization of it,” said Harrison. 
That transition also means internal developers are expanding their skill sets to include more capabilities, enhancing their resumes and growing their careers while also contributing to the growth of the company itself. It’s a win-win situation. “They know that the logs are collected in Splunk and through VictorOps, they get that notification back when something’s gone wrong. It’s just part of how they do their daily job now,” said Harrison. 
How did they manage to put all this power into the hands of the developers themselves? “Early adoption of OpenShift 4 was critical to our success. We were very much looking at this as if we were going to be on this for the next five years and didn’t want to get stuck on an older version, so we took that risk. We were very tightly working  with Red Hat while doing this. They’ve been embedded with our teams for the last year. They are a part of our team. We drove innovation within our team and within Red Hat. The stuff we were working on was being brought back into Red Hat to bring the platform to where it is today,” said Harrison.
“That transformation we talked about earlier from traditional systems administrators in an operational role to devops engineers, we did that in less than a year. We weren’t doing this a year ago. We were traditional systems administrators in our silos. We had Linux guys, we had VMware guys, we had network guys. Now we’re all on a team together, we’re all doing this stuff every day. It’s been an amazing transformation for the technology we’re working in and for our careers. It’s been great, and that tight partnership facilitated that shift very easily. Having Operators available to us allowed us to deliver services to the dev teams almost immediately on day one,” said Harrison.
 
The post How Omnitracs Transformed to a DevOps Culture with OpenShift appeared first on Red Hat OpenShift Blog.
Quelle: OpenShift