Introduction
Somewhere inside a global agriscience’s organization, a scientist is trying to answer a simple question: has this active ingredient been used in any other project, under any other name? For years, the honest answer was that nobody could say for certain. That uncertainty was the shape of a much larger problem, and the stakes behind it were bigger than any one team’s workflow.
The agrisciences organization Syngenta, plays a critical role in helping growers protect crops from weeds, disease, and pests while supporting global food security. Yet bringing a new product from concept to market already requires R&D cycles that can span eight to twelve years. In this environment, data friction is more than an operational challenge. Every month lost to fragmented data, inefficient workflows, or delayed insights extends time-to-market and delays growers’ access to safer, more effective products. Reducing this friction across the R&D lifecycle is therefore not simply about improving productivity; it is a strategic lever for accelerating innovation and impact.
Yet critical business data was locked in over disparate R&D data sources. The organization needed a scalable way to ingest, standardize, and productize this data to support analytics, machine learning, and cross-functional decision-making, while also managing high-volume streaming data from IoT-connected labs generating billions of data points. Without a unified platform, teams across its five domains: Portfolio, Research, Product Safety & Regulatory, Product Design, and Real-World Data, operated in silos, limiting data reuse, slowing innovation, and increasing the cost and risk of duplicated data engineering effort.
For Syngenta, solving this wasn’t a matter of adding more people or building one more integration on top of the ones that already existed. What the organization needed was a different kind of fix entirely: a platform where standardization and ownership were built into the architecture itself, so that reuse became the default outcome of building something, not a follow-up project someone had to remember to do later.
This case study covers Modak’s multi-year partnership with Syngenta to design and build, a cloud-based enterprise data platform (Gaia) now producing 82 reusable of reusable data products across all five R&D domains. The result is an organization where data that once took months to locate and trust is now standardized, governed, and reusable by design, freeing their scientists and decisionmakers to spend less time reconciling data and more time getting safer, more sustainable crop protection solutions to the growers who depend on them.
Before Modak: The Challenge
Across the organization’s five R&D domains: Portfolio, Research, Product Safety & Regulatory, Product Design, and Real-World Data, each team had spent years building exactly the systems and workflows their own function needed. That specialization worked well in isolation, but it came at a real cost when work needed to cross domain lines.
A researcher trying to connect a lab-finding to a regulatory submission, or a portfolio manager trying to trace an ingredient’s full project history, was doing manual detective work before any actual analysis could begin. The deeper cost sat one layer below that daily friction: there was no shared definition of what counted as a trustworthy, reusable dataset. Five teams could each independently reach the same conclusion about the same underlying data and produce five different, non-comparable versions of it, with no way for any of them to know the others existed. Effort that should have compounded across the organization instead reset to zero every time a new team touched a familiar problem.
- No common standard for what made a dataset trustworthy or reusable across domains
- Manual, ad hoc integration work required for any question that crossed a domain boundary
- The same underlying data problems solved independently, and differently, by multiple teams
- No organizational visibility into how much of that effort was avoidable duplication
Technically, the scale and shape of the data made this a genuinely hard engineering problem as well. More than 80 disparate source systems, spanning thousands of tables, held everything from structured lab results to regulatory documents to field trial records, each in its own format, on its own schedule, using its own identifier conventions for what were often, underneath it all, the same real-world entities. There was no standardized way to ingest, model, and productize this data consistently, so every new use case started from raw material rather than from a trusted foundation someone could build on.
That structural sprawl was compounded by a velocity problem legacy R&D infrastructure had never been built to handle. IoT-connected lab instruments alone were generating over 13 billion plus data points, demanding real-time streaming ingestion running in parallel with conventional batch analytics rather than a separate system bolted afterward. Getting both to run reliably on one foundation, without either one compromising the other, was the technical prerequisite for everything else on the platform depended on.
- 80+ disparate sources and thousands of tables, each with its own format and identifier scheme
- No consistent method for turning raw source data into standardized, analysis-ready products
- 13B+ IoT data points requiring streaming ingestion, not just scheduled batch loads
- Batch and streaming workloads needing to coexist reliably on shared infrastructure, not run as separate systems
Modak’s Approach
Modak’s response to this problem started with a decision that shaped everything downstream: standardization and ownership needed to be built into the architecture itself, not layered on as governance policy after the fact. Rather than adding another team to manage integrations centrally, Modak worked with Syngenta to design the Gaia platform, where reuse was the natural output of doing the work once, correctly, rather than a discipline someone had to enforce after the fact.
The new Gaia data platform was built on Databricks and AWS, with Unity Catalog providing governance, lineage, and access control across every layer, and Delta Lake serving as the reliable storage foundation beneath it. Choosing Databricks as the foundation gave the organization a single environment capable of handling both the batch-oriented analytical workloads most R&D data still run on, and the real-time streaming demands of IoT-connected lab telemetry, without forcing a trade-off between the two.
- Standardization and domain ownership designed into the platform architecture, not enforced through policy afterward
- Databricks and AWS as the unified foundation for both batch and streaming workloads
- Unity Catalog governing lineage, access, and discoverability across all five domains from day one
- Delta Lake as the reliable storage layer underpinning every data product built on top of it

Ingestion was designed around the organization’s actual operating rhythm rather than a generic batch schedule: intraday pipelines refreshing approximately every two hours run continuously alongside real-time streaming ingestion from IoT-connected labs, so both scheduled analytical work and near-real-time telemetry live on the same platform instead of two disconnected systems maintained by two different teams.
Reliability and cost discipline were treated as architectural requirements, not operational afterthoughts. A purpose-built orchestration and observability auxiliary layer give the platform lineage, logging, and pipeline transparency end to end, while CI/CD pipelines and DevOps practices keep deployments consistent and scalable as the platform grows. FinOps principles were built into workload design from the outset, so cost accountability scaled alongside the platform rather than becoming a retrofit years into its life.
- Intraday batch ingestion (~2hr cycles) running in parallel with real-time streaming ingestion, not as separate systems
- Forge framework delivering platform-wide lineage, logging, and observability
- CI/CD and DevOps practices supporting reliable, repeatable deployment at scale
- FinOps principles built into workload and pipeline design from the start, not added later
Building the Platform: Engineering Reuse into the Architecture
The organization’s five R&D domains didn’t need five different platforms adapted to their individual quirks. They needed one platform disciplined enough to serve all five without diluting what made any of them distinct, and that discipline had to be engineered in from the first pipeline, not negotiated domain by domain after the fact.
Every source, regardless of which domain it belonged to, moved through the same progression: raw data preserved exactly as it arrived, standardized into one agreed definition per entity, and only then published as a governed, purpose-built data product.
- Spark processed structured business records, scientific data formats, and streaming telemetry as citizens of the same system, no separate tool for genomics, no bolt-on for IoT
- Ingestion ran on two clocks at once: intraday batch pipelines and continuous streaming, so neither core analytics nor high-velocity lab data was sacrificed for the other
Governance was not a compliance layer added after the platform worked. Every data product carries lineage back to its source by default, access is governed at the level of precision that regulated environments actually require, and any product built by any domain is discoverable by every other domain, rather than known only to the team that happened to build it. That combination: provable lineage + cross-domain visibility, is what let Product Safety & Regulatory meet traceability standards without a parallel audit process.
And lastly, but arguably the hardest: making all of this usable by people who were not data engineers. Scientists, project managers, and coordinators needed direct, governed access to trusted data without routing every question through a technical intermediary. Meeting that bar is what turned the platform from infrastructure into something the organization’s day-to-day decision-making actually runs on.
The Five Domains, Built on One Foundation

What distinguishes this implementation from a typical platform migration: the five domains did not receive five customized solutions. They received the same foundation, and each proved it could carry genuinely different scientific and operational weight without buckling.
| Domain | No of Sources | No of Products | What Was Built | Technologies Used |
| Portfolio | 11 | 13 | An ingredient identity resolution product mapping every known name and identifier to one canonical entity; a consolidated project summary product combining milestones, cost, and projected value into a single record per project; a stage-gate tracking product monitoring pipeline velocity across the project lifecycle; an environmental impact scoring pipeline calculating sustainability metrics at both ingredient and product level; a cost-estimation product built on market pricing data | Databricks SQL, Delta Lake, Unity Catalog |
| Research | 14 source systems | 14 in prod | A large-scale research results product consolidating tens of thousands of research runs annually from core crop protection science; genomics and mode-of-action data products covering gene-protein relationships and expression data; a decision-support product combining research and portfolio data to inform project prioritization; a shared reference-data layer providing standardized master lookups reused across every other Research product | Spark, Delta Lake, streaming ingestion, Unity Catalog |
| Product Safety & Regulatory | 14 | 15 in prod, 2 in dev | A study and data management product tracking commissioning, results, and metadata across toxicology, eco-toxicology, and exposure studies; regulatory dossier content structured and aligned to formal submission templates; a master regulatory dataset directly powering submission workflows, safety data waivers, and early document-assembly automation | Unity Catalog (lineage, access control), Delta Lake |
| Product Design | 7 in prod, 3 in progress | 15 in prod | An organizational data product providing a current, automatically refreshed view of personnel and team structure; a project and task tracking product connecting individual work items to the broader design pipeline; a formulation compatibility product covering tank-mix assessment results; a time-tracking product aggregating effort across projects and activity categories | Databricks SQL, automated Databricks workflows |
| Real World Data | 25 | 16 in prod | A flagship curated trial-results product harmonizing over a decade of current and legacy field trial data without altering original records; drone-based aerial assessment products; real-time IoT field telemetry products covering environmental conditions; protocol design and trial planning products; trial-station and operational data products | Spark, Delta Lake, Trial Data Model (star schema), streaming + batch ingestion |
Results That Compound: Platform and Domain Impact
What started as an effort to make R&D data usable across five disconnected domains became something larger: a governed platform that the agriscience organization now treats as core infrastructure, not a project with a defined end date. The results show up at two levels: the scale and reliability of the platform itself, and the specific ground each domain gained by building on it.
Platform-Level Impact
- 80+ R&D data sources and thousands of tables unified onto a single governed foundation
- 82 reusable data products now live across the five domains, each built once and queried by every team that needs it
- ~10 billion records processed monthly, scaling to `170 billion records processed annually, sustained at a ~99%+ job success rate
- 13B+ billion IoT data points streamed continuously from connected lab instruments, integrated with batch-ingested data rather than isolated in a separate system
- A cross-domain hackathon spanning five use cases, with teams building new solutions in days on data they didn’t know existed the week before
The more consequential shift is harder to reduce to a stat: standardization and ownership, once something someone had to enforce, became the default architecture teams build within. A cross-domain question that once triggered weeks of manual reconciliation now starts from a trusted, governed product someone else already built.
Domain-Level Impact
| Domain | Before Modak | After Implementation |
| Portfolio | Ingredient identity untraceable across systems, exposing R&D investment decisions to blind spots and duplicated research spend | Ingredient identity resolved platform-wide via a governed lookup product, giving investment and portfolio decisions a single source of truth |
| Research | Research output trapped in disconnected systems, slowing the path from scientific result to regulatory or commercial decision | Tens of thousands of research runs annually flow into standardized, governed products, cutting the time from result to decision and putting scientific and business data on equal footing |
| Product Safety & Regulatory | Traceability assembled manually per submission, adding time and audit risk to a process regulators hold to exacting standards | Lineage and traceability built into every data product by design, shortening submission cycles and reducing audit exposure |
| Product Design | Operational visibility gated behind engineering requests, slowing day-to-day decisions for project and resourcing teams | Data products consolidating material, chemical, and formulation reference data; laboratory testing and process data; project and task tracking; and time-tracking across projects, giving teams across PTE a single, standardized view of product design data instead of siloed, per-function versions. |
| Real World Data | Field performance evidence fragmented across current and legacy systems, delaying the yield and efficacy analysis that validates R&D investment | A decade-plus of trial data harmonized into one trusted, analysis-ready foundation, accelerating the evidence base behind product and yield decisions |
Looking Ahead
Syngenta’s new Gaia data platform now supports a growing base of self-service users across all five domains, where scientists, project managers, and analysts who once depended on engineering teams to access and interpret data, and who now work with it directly as part of their daily decision-making. That shift in who touches data, and how confidently, is arguably the more durable outcome of this work, more so than any single metric.
Across the agrisciences industry, the pressure this kind of platform responds to isn’t unique to any one organization. R&D cycles remain long, regulatory scrutiny remains high, and the cost of duplicated data engineering effort compounds the same way industry-wide. What’s changed here is that decisions once settled by whoever had spent the most time with a particular dataset are increasingly settled by teams working from the same governed, trusted foundation, reducing the friction between having a question and getting a defensible answer.
Syngenta now continues to expand where this foundation is used, including potential applications of AI and other capabilities that could further support how teams work with data and the platform. Standardization, governance, and ownership are built into the architecture rather than layered on after the fact, creating a consistent foundation for how data is managed and consumed. For an industry racing to deliver safer, more sustainable crop protection solutions to growers faster, that discipline is no longer simply a data engineering detail. It is becoming part of how the work gets done.


