Skip to main content

Big Data for Agri-Food: Principles and Tools

WUR

About This Course

Demystify complex big data technologies

Compared to traditional data processing, modern tools can be complex to grasp. Before we can use these tools effectively, we need to know how to handle big data sets. You will understand how and why certain principles – such as immutability and pure functions – enable parallel data processing (‘divide and conquer’), which is necessary to manage big data.

During this course you will acquire this principal foundation from which to move forward. Namely, how to recognise and put into practice the scalable solution that’s right for your situation.

The insights and tools of this course are regardless of programming language, but user-friendly examples are provided in Python, Hadoop HDFS and Apache Spark. Although these principles can also be applied to other sectors, we will use examples from the agri-food sector.

Data collection and processing in an Agri-food context

Agri-food deserves special focus when it comes to choosing robust data management technologies due to its inherent variability and uncertainty. Wageningen University & Research’s knowledge domain is healthy food and the living environment. That makes our data experts especially equipped to forge the bridge between the agri-food business on the one hand and data science and artificial intelligence (AI) on the other.

Combining data from the latest sensing technologies with machine learning/deep learning methodologies allows us to unlock insights we didn’t have access to before. In the areas of smart farming and precision agriculture this allows us to:

  • Better manage dairy cattle by combining animal-level data on behaviour, health and feed with milk production and composition from milking machines.
  • Reduce the amount of fertilisers (nitrogen), pesticides (chemicals) and water used on crops by monitoring individual plants with a robot or drone.
  • More accurately predict crop yields on a continental scale by combining current with historic data on soil, weather patterns and crop yields.

In short, this course’s foundational knowledge and skills for big data prepare you for the next step: to find more effective and scalable solutions for smarter, innovative insights.

Is this course for you?

You are a manager or researcher with a big data set on your hands, perhaps considering investing in big data tools. You’ve done some programming before, but your skills are a bit rusty. You want to learn how to effectively and efficiently manage very large datasets. This course will enable you to see and evaluate opportunities for the application of big data technologies within your domain. Enrol now.

This course has been partially supported by the European Union Horizon 2020 Research and Innovation program (Grant #810775, “Dragon”).

What You Will Learn

After successful completion of this course, you will be able to:

  • Understand key big data characteristics and principles such as volume, velocity, variety, veracity, immutability and pure functions.
  • Distinguish between scaling up and scaling out, and know when to apply each.
  • Process big data using map-reduce, clusters and distributed file systems like Hadoop.
  • Work efficiently with dataframes, wrapper technologies and datalakes using lazy evaluation.
  • Apply big data workflows and pipelines to your own case or organisation.

Grow these skills

Big Data Fundamentals
Big Data Problem Assessment
Scalable Data Processing
Distributed Computing
Hadoop Distributed File System (HDFS)
Apache Spark Data Processing
MapReduce Programming Concepts
Data Pipeline Design
Data Lake Architecture
Big Data Workflow Implementation

Curriculum

6 Weeks, 6-10 hours per week


Module 1: Big data definition and characteristics
Learn how to recognize the characteristics of a big data problem in agriculture and identify where the biggest challenge lies. Explore whether the solution should focus on data volume, velocity, variety or veracity, and understand when to scale up or scale out.

Module 2: Big data principles: what are they and why do we need them
Discover the core principles required for scaling out, including immutability, pure functions and the map-reduce paradigm. Learn what these concepts are and why they are essential for efficient big data processing.

Module 3: Bring those principles to practice
Put big data principles into practice by learning how clusters and distributed file systems work. Explore Hadoop, client-server architecture and understand why distributed systems provide scalable solutions for processing large datasets.

Module 4: Big data technologies that make implementation so much easier
Explore the big data technology stack and discover how modern tools simplify implementation. Learn how platforms such as Apache Spark automatically apply map-reduce concepts, making large-scale data processing more efficient.

Module 5: The big data workflow and pipeline; the how and why of datalakes
Dive deeper into data management by exploring datalakes and understanding how they differ from traditional databases. Learn what a big data workflow looks like, how data pipelines are structured, and why these approaches support scalable data processing.

Requirements

A university education and/or working knowledge of math and science and, of course, being a computer science enthusiast will help a lot!

Meet the instructors

Ioannis N. Athanasiadis

Professor in Artificial Intelligence and Data Science at Wageningen University & Research.

Sjoukje Osinga

Assistant Professor in Information Technology at Wageningen University & Research.

Christos Pylianidis

PhD Student in Information Technology at Wageningen University & Research.

Course Summary

  1. Course Number

    AIN90010
  2. Classes Start

  3. Classes End

Enroll