View a markdown version of this page

LS-S01 Research and discovery - Life Sciences Lens

LS-S01 Research and discovery

Research and discovery (R&D) represents the crucial first phase in life sciences innovation, where organizations work to identify and develop potential drug candidates. This phase involves complex data analysis across multiple domains including drug discovery, bioinformatics, genomics, and laboratory operations. Organizations face challenges in managing and analyzing vast amounts of complex scientific data, implementing efficient computational workflows, and maintaining regulatory adherence while accelerating the pace of discovery. Modern approaches use advanced technologies like AI/ML, high-performance computing, and cloud storage solutions to process and analyze data at scale, while adhering to Good Laboratory Practices (GLP) and Findable, Accessible, Interoperable, and Reusable (FAIR) data principles.

Use cases

  • LS-S01-UC01 Drug discovery: Drug discovery is a costly and lengthy process. Pharmaceutical companies are actively seeking to accelerate the discovery phase, especially in areas such as understanding of disease mechanisms using multi-modal data analysis and rapidly identifying molecular targets and candidate drug leads. These steps involve collection, standardization, storage, management, and governance of large amounts of complex data along with the availability of specialized compute to generate insights from them. With the advent of specialized AI models, pharmaceutical organizations are seeking to use them to advance their drug discovery timelines by orders of magnitude in a secure, transparent, and ethical manner.

  • LS-S01-UC02 Bioinformatics: Bioinformatics plays a key role in providing as well as using the necessary tools and databases that enable researchers and clinicians to interpret biological data. Bioinformaticians and data scientists, in most types of organizations, such as clinical diagnostics, population sequencing, pharmaceutical R&D, and agriculture, have the necessary skills to build and use pipelines that process raw data, such as genomics data, and interpret that data using analytics tools, machine learning models, and statistical analysis. Bioinformatics teams seek to simplify their analysis lifecycle by focusing more on the science and less on infrastructure. Additionally, they constantly seek to optimize the performance, accuracy, and cost of analysis for better scientific and business outcomes.

  • LS-S01-UC03 Genomics: Genomics and other omics data (like transcriptomics or proteomics) play a crucial role in clinical diagnostics, drug discovery, gene editing, precision medicine, and bio-surveillance. Omics data typically has a large footprint (for example, raw and processed genomics data for single human patient can range from 100 to 200 GB). Based on the type of project, organizations can maintain tens of thousands, sometimes millions of patients worth of data. Such organizations face challenges to manage the cost and governance of this data. They actively seek to simplify data lifecycle, implement FAIR principles within and across organizations, and reduce short and long-term storage and usage costs. Raw omics data needs to be processed using bioinformatics workflows to derive insights from them. Organizations actively look to process these data at scale in the most performant and cost-efficient manner. They are faced with several options to do so, whether running it themselves in a datacenter on-premises, using a vendor that provides a pre-built experience, or building their own infrastructure in the cloud. They seek guidance on the best option for them given their goals, team size, internal skillsets, budget, and existing environmental constraints.

  • LS-S02-UC04 Connected labs: R&D organizations are seeking to modernize their laboratory infrastructure, aiming to use data to enhance research capabilities. A key focus of this transformation is optimizing laboratory data integration and accessibility. By addressing the current state of siloed data across laboratory instruments, Laboratory Information Management Systems (LIMS), and Electronic Lab Notebooks (ELN), companies seek to enable data integration between on-premises data and cloud analytics while following FAIR data principles. This forward-thinking approach empowers scientists with more comprehensive and readily available data, accelerating insights and innovation. Furthermore, the establishment of robust data governance for unstructured lab data improves the usage of AI/ML and advanced analytics to accelerate drug discovery.