Home / Resources / Blogs

Agile in Action — Big Data Projects

Technical JUL 5, 2018 LUMIQ Team Data Engineering

Big data has been ‘in’ and booming for the last few years. While some organizations still maintain a level of skepticism, others are quick to jump onboard and transition to big data too quickly. However, in both cases, it’s important to understand how big data can help businesses innovate and grow.

The first thing to understand is the nature of big data. Big data projects are usually of an uncertain nature. Hence, it is crucial to be able to rely on a delivery method that allows the project to consume changing requirements or quickly change directions altogether.

This is where ‘Agile’ methodology comes in.

When I say Agile, I’m referring to Scrum Methodology, as opposed to other Agile variations such as AUP, DSDM, Adaptive Software Development and more.

Businesses who already favor Agile do so because:

  • The basic framework is built on the principle of work breakdown. If we drill down requirements to the most granular level, it makes it faster and simpler to develop, test and deploy.
  • The second most important factor is getting business value for every sprint. This is where ‘Agile’ methodology comes in.
  • It also allows projects to learn from the feedback, take in new or changing requirements and quickly change direction when necessary, without changing the process at all.

Bringing Agile to Big Data

Agile is well suited for traditional application development projects, but the most data-intensive projects (such as data warehousing, business intelligence and big data) face particular challenges across the entire project lifecycle, from requirements to UAT. Instead of striving for perfection, Agile aims for the minimum viable product (MVP), which means that agile produces concrete results faster than traditional development methods. The whole point of agile is finding out quickly whether your idea works or not. Agile embodies the concept of “failing fast,” a phrase that sums up the ethos of modern innovation.

In such projects, we have large amounts of unstructured data, which we want to use to our advantage. And this is where big data projects with their abstract requirements can benefit the most from Agile methodology. Abstractions are tolerable when you have plenty of time and money, but when time and money are tight, you need concrete results quickly. The whole point of Agile is finding out quickly whether your idea works or not.

Challenges & Best Practices while using Agile for Big Data Projects

1:One of the greatest challenges I have encountered in implementing an Agile approach for big data projects is the work culture. Agility, especially in the context of big data, takes a lot of effort and really hard work. Creating an Agile environment doesn’t mean turning off governance, eliminating documentation, and giving developers a sandbox to work in.

All stakeholders must understand and appreciate the fundamentals of Agile and be open to the fact that, as with every other process, this one requires customization. One example is the sprint duration. Depending on user stories and effort estimation, we could plan for a three or four-weeks sprint instead of a two-week sprint.

2 In most projects, Scrum teams do not have all prerequisite activities and tasks planned and executed in order to begin development in sprint one. As I have written above, since big data projects inherently have abstract requirements and large volumes (running in terabytes) of data, this creates added challenges.

It’s important to be ready with the environment setup, data architecture, product backlog with prioritized user stories, a sufficient number of sprint-ready user stories, etc. so that the Agile approach for big data projects can be implemented. This helps teams avoid any roadblocks in the first few sprints, which is when the customer looks to validate the approach and capability of the Scrum team.

3. Training the Business

Agile is not some kind of free-form anarchy — it’s a structured way of producing usable software without having a pre-written script. Thus, it embraces both discipline and flexibility.

Explaining Agile to people working in the business units is as important as training the IT people. Business people need to understand what the IT team is doing. Here, co-locating the IT and Business team helps in collaboration and communication during project execution.

In terms of access to the management of big data projects, it is appropriate to iterate in the following steps: defining the problem, identifying the knowledge or information gap, setting hypothesis, test hypotheses (accept or reject), continue in the cycle (iteration).

How to Begin

Companies looking for ways to bring Agile into their big data operations should begin with simple steps. This includes outlining key use cases and forming developer teams with diverse backgrounds. Big data projects need to involve a person such as a manager or an analyst who has the domain knowledge and is able to ask the right questions.

Conclusion

Big data is a lucrative domain to tap into. But, when it comes to the implementation, what approach should be applied in the management of big data projects? If we consider the comparison based on the triple constraint of the project, we can determine that for the management of big data projects, Agile is the preferable approach. This is because of the constantly changing requirements and uncertainty that comes with it.

As stated above, in implementing big data projects, it is recommended to start with small use cases, accepting small failures that move the team forward and continue the iterative approach. Slowly but surely — just as the Agile flavored project — the team will become well versed in handling big data.

Ready to turn your data into decisions?

Tell us where your data is slowing you down. We will show you what production-grade looks like in your own AWS cloud.

Book a briefing