BIOGRAPHICAL SKETCH
Provide the following information for the Senior/key personnel and other significant contributors.
Follow this format for each person. DO NOT EXCEED FIVE PAGES.
NAME: Mattias Johansson
eRA COMMONS USER NAME (credential, e.g., agency login): M.JOHANSSON
POSITION TITLE: Scientist
EDUCATION/TRAINING (Begin with baccalaureate or other initial professional education, such as nursing,
include postdoctoral training and residency training if applicable. Add/delete rows as necessary.)
Completion
DEGREE
Date FIELD OF STUDY
INSTITUTION AND LOCATION (if applicable)
MM/YYYY
Master of Science in Engineering
Umeå University, Umeå, Sweden 06/ 2003
Physics, Mathematical Statistics
Epidemiology and genetics of
Umeå University, Umeå, Sweden PhD 06/ 2008
prostate cancer
International Agency for Research on Postdoctoral Epidemiology and genetics of
11/ 2010
Cancer (IARC) training smoking-related cancers
A. Personal Statement
I am molecular epidemiologist working in the Genomic Epidemiology Branch at the International Agency for
Research on Cancer (IARC/WHO). Together with Dr Robbins, I lead the Risk Assessment and Early Detection
(RED) team which includes 16 scientists, postdocs and support staff. We have a broad research agenda
focusing on multiple cancer sites with two overarching aims; i) to elucidate etiological and mechanistic factors
contributing to cancer risk and outcome, and ii) to develop accurate prediction models for early detection and
prognostics. A major theme throughout these studies is integrating information from questionnaire,
demographic, biomarker, genetic and genomic data. These studies are typically conducted through
international collaborative efforts, and I have extensive experience in coordinating large consortium projects,
including the HPV Cancer Cohort Consortium (HPVC3) and the Lung Cancer Cohort Consortium (LC3) which
involves 26 cohorts from around the world. Recently, my team has focused on developing and validating tools
for cancer risk prediction and early diagnosis (Muller et al. BMJ. 2019). I am an MPI of the INTEGRAL-AT
program (U19) where I co-lead Project 2 together with Dr Robbins. We have published several papers on
cancer risk prediction, screening, and biomarkers that are relevant to the current proposal (e.g. Guida et al.
JAMA Onc. 2018, Robbins et al. Br J Cancer. 2021). Most recently, I led a large-scale discovery study of pre-
diagnostic protein biomarkers for lung cancer (LC3, Nat. Comm. 2023, Feng et al. JNCI. 2023).
Ongoing projects that I would like to highlight include:
NIH 2U19CA203654-07 (Amos, Hung Johansson) 08/01/23-03/31/28
National Institute of Health (NIH)
Integrative analysis of lung cancer etiology and risk – Application and Translation (INTEGRAL-AT) - Project 2
component: Integrating Biomarkers into Lung Cancer Risk Profiling
The main goals of Project 2 are to develop a biomarker-based risk model, evaluate its acceptability in a real-world
screening setting, and identify risk markers for never smoking lung cancer.
Role: m-PI, Project leader
NIH 7U19CA203654-02 (Amos, Hung Johansson) 08/01/17-03/31/24
National Institute of Health (NIH)
Integrative analysis of lung cancer etiology and risk - Project 2 component: Biomarkers of lung cancer risk
The main goal is to pursue a comprehensive and complementary set of research agendas to understand predictors
of smoking and lung cancer.
Role: m-PI, Project leader
CRUK C18281/A29019 (PI: Martin) 10/01/20-09/30/25
Cancer Research UK
Reducing the Burden of Cancer: causal risk factors, mechanistic targets and predictive biomarkers
The major goals are to provide a robust epidemiological evidence-base for the future development of lifestyle and
nutritional interventions in people at risk of, or diagnosed with, cancer; and to identify novel biomarkers of cancer
risk and progression.
Role: co-PI
B. Positions and Honors
2001 - 2002 Program management, Department of Physics, Umeå University, Sweden
2004 Research engineer, Akzo Nobel Surface Chemistry, Örnsköldsvik, Sweden
2004 - 2005 Research engineer, Department of Urology, Umeå University, Sweden
2010 - 2012 Temporary staff scientist, Genetic Epidemiology Group, International Agency for Research on
Cancer (IARC), Lyon, France
2012 - Fixed-term staff scientist, Genetic Epidemiology Group, International Agency for Research on
Cancer (IARC), Lyon, France
C. Contributions to Science
Link to all publications: https://www.ncbi.nlm.nih.gov/myncbi/mattias.johansson.1/bibliography/public/
Projects related to risk prediction and development of biomarkers for lung cancer
In anticipation of widespread low-dose CT screening for early detection of lung cancer, we initiated a major
research program focusing on lung cancer risk modelling and identifying biomarkers that are useful in
identifying individuals who are likely to benefit from screening. The initial INTEGRAL U19 grant where I lead
the biomarker component as m-PI (Project 2) supports most of this work. Our findings initially demonstrated
that that circulating tumor-related protein biomarkers have a strong potential to identify those individual
destined to develop lung cancer, over and above risk information that can be gained from traditional smoking-
based risk prediction models (Guida et al. JAMA Oncology. 2018). In parallel, we have also evaluated the
benefits and harms of lung cancer screening using the NSLT data (Robbins et al. Lancet Respir. Med. 2019).
We subsequently carried out a major discovery analyses for early protein markers within the LC3 consortium
and identified 36 individual protein biomarkers that predispose lung cancer diagnosis, most of which were
novel and improve lung cancer risk discrimination (LC3. Nat. Comm. 2023, Feng et al. JNCI. 2023).
• Guida F, Sun N, Bantis LE, … Johansson M,* Hanash S*. Assessment of Lung Cancer Risk on the Basis
of a Biomarker Panel of Circulating Proteins. JAMA Oncol. 2018. PMID: 30003238 *Senior author
• Robbins HA, Callister M, Sasieni P ,…, Johansson M.* Benefits and harms in the National Lung Screening
Trial: expected outcomes with a modern management protocol. Lancet Respir Med. 2019. PMID:
31076382 * Senior author
• The Lung Cancer Cohort Consortium (LC3). The Blood Proteome of Imminent Lung Cancer Diagnosis.
Nat. Com. 2023 Aug 1. Senior author
• Feng X, Wu WY, Onwuka JU, …, Johansson M. Lung cancer risk discrimination of prediagnostic
proteomics measurements compared with existing prediction tools. J Natl Cancer Inst. 2023. PMID:
37260165; PMCID: PMC10483263. Senior author
Biomarker studies on cancer etiology
A particular focus spanning my entire research career involves circulating biomarkers of the one-carbon
metabolism, such as folate and vitamin B12, initially in relation to prostate cancer development, and more
recently in relation to lung, head and neck, and kidney cancer. Whilst the initial studies on prostate cancer did
not provide any clear support for an important role of the one-carbon metabolism pathway in prostate cancer
etiology (Johansson et al. CEBP. 2008), our initial study on lung cancer suggested strong inverse associations
of several one-carbon metabolism biomarkers with lung cancer risk independently of smoking status, in
particular vitamin B6, folate, and methionine (Johansson et al. JAMA 2010). This analysis triggered the
formation of the Lung Cancer Cohort Consortium (LC3) where we concluded that folate, vitamin B6 and
methionine were unlikely to substantially modify lung cancer risk (Fanidi et al. JNCI 2018). The LC3
collaboration has generated multiple additional biomarker studies, such as our evaluation of C-reactive protein
in lung cancer (Muller et al. BMJ 2019).
• Johansson M, Appleby PN, Allen NE et al. Circulating concentrations of folate and vitamin B12 in relation
to prostate cancer risk: results from the European Prospective Investigation into Cancer and Nutrition
study. Cancer Epidemiol Biomarkers Prev. 2008. PMID: 18268110.
• Johansson M, Relton C, Ueland PM, et al. Serum B vitamin levels and risk of lung cancer. JAMA. 2010.
PMID: 20551408.
• Fanidi A, Muller, DC, Yuan, JM, …, Johansson M.* Circulating Folate, Vitamin B6, and Methionine in
Relation to Lung Cancer Risk in the Lung Cancer Cohort Consortium (LC3). J Natl Cancer Inst. 2018.
PMID: 28922778 *Senior author
• Muller DC, Larose TL, Hodge A, …, Johansson M.* Circulating high sensitivity C reactive protein
concentrations and risk of lung cancer: nested case-control study within Lung Cancer Cohort Consortium.
BMJ. PMID: 30606716 *Senior author
Projects related to risk prediction and development of biomarkers for head and neck cancer
Our initial analysis on antibodies against human papillomavirus (HPV) in head and neck cancer demonstrated
that antibodies against HPV16 E6 can be detected for 35% of future oropharynx cancers, a similar proportion
of the underlying fraction of HPV driven oropharynx cancer in Europe, in turn suggesting that this represents
the most HPV driven oropharynx cancers (Kreimer & Johansson et al. JCO. 2013). That the HPV16 E6
antibodies were only present in a handful of controls (0.6%), and that the sensitivity to predict future
oropharynx cancers were constant during the whole follow-up period, suggest HPV16 E6 to be an extremely
promising early cancer marker with unparalleled of sensitivity of above 80% and specificity of over 99%. We
subsequently evaluated the biomarker in relation anogenital cancers and demonstrated that it is much less
sensitive to predict other HPV related cancers, but that they would need to be considered in follow-up of
HPV16 E6 healthy subjects (Kreimer et al. J Clin Oncol. 2015). These studies led to developing the HPV
Cancer Cohort Consortium (HPVC3) which I coordinated to evaluate the potential of HPV16 E6 as screening
tool (Kreimer et al. Ann. Oncol. 2019). We recently published the final risk model for HPV16 E6-based
prediction of oropharyngeal cancer (Robbins et al. JCO 2022) which highlighted that HPV16E6 positive men
have substantial absolute risk of oropharyngeal cancer.
• Kreimer AR*, Johansson M*, Waterboer T, et al. Evaluation of human papillomavirus antibodies and risk
of subsequent head and neck cancer. J Clin Oncol. 2013. PMID: 23775966 *Joint first-authors
• Kreimer AR, Brennan P, Lang Kuhs K, … Johansson M*. Human papillomavirus antibodies and future risk
of anogenital cancer: a nested case-control study in the European Prospective Investigation into Cancer
and Nutrition (EPIC) study. J Clin Oncol. 2015. PMID: 25667279 *Senior author
• Kreimer AR, Ferreiro-Iglesias A, Nygard M, …, Johansson M.* Timing of HPV16-E6 antibody
seroconversion before OPSCC: findings from the HPVC3 consortium. Ann Oncol. 2019. PMID: 31185496
*Senior author
• Robbins HA, Ferreiro-Iglesias A, Waterboer T, …Johansson M,* … Absolute Risk of Oropharyngeal
Cancer After an HPV16-E6 Serology Test and Potential Implications for Screening: Results From the
Human Papillomavirus Cancer Cohort Consortium. J Clin Oncol. 2022. PMID: 35700419 *Senior author
Studies on renal cancer
Since moving to IARC in 2008, together Dr Brennan, I co-lead a collaboration with colleagues at the US NCI
(Dr Chanock and Dr Purdue) on genome-wide association studies (GWAS) on renal cancer. This led to the first
GWAS publication on renal cancer that identified three susceptibility loci on 2p21 (the EPAS1 gene), 11q13.3,
and 12q24.31 (the SCARB1 gene) (Purdue & Johansson et al. Nat Genet 2011). Through subsequent
collaborations we have expanded the study population, and now await the third iteration of the renal cancer
GWAS with over 20,000 cases. In parallel, I have led various biomarker studies on renal cancer, the most
relevant being an evaluation of pathways related to one-carbon metabolism that demonstrated that individuals
with higher levels of pre-diagnostic circulating pyridoxal 5’-phosphate (PLP) have substantially increased renal
cancer risk, as well as poorer survival following diagnosis (Johansson et al. JNCI 2014). We have also carried
out a series of studies employing Mendelian randomization techniques to evaluate the causal relevance of
putative risk factors. Amongst these, perhaps the most informative was that on renal cancer and obesity
related risk factors that firstly confirmed elevated body-mass index as an important cause of renal, and
secondly, identified elevated diastolic blood pressure and circulating insulin as having a novel and important
role in mediating the established relation between obesity and renal cancer risk (Johansson et al. PLoS Med.
2019). Most recently we published results from the MetKid consortium that highlighted a major role of the blood
metabolome in renal cancer development (Guida et al. PLoS Med. 2021).
• Purdue MP*, Johansson M*, Zelenika D, et al. Genome-wide association study of renal cell carcinoma
identifies two susceptibility loci on 2p21 and 11q13.3. Nat Genet. 2011. PubMed PMID: 21131975; *Joint
first-authors
• Johansson M, Fanidi A, Muller DC, et al. Circulating biomarkers of one-carbon metabolism in relation to
renal cell carcinoma incidence and survival. J Natl Cancer Inst. 2014. PMID: 25376861
• Johansson M, Carreras-Torres R, Scelo G, et al. The influence of obesity-related factors in the etiology of
renal cell carcinoma-A mendelian randomization study. PLoS Med. 2019. PMID:30605491
• Guida F, Tan TY, Corbin LJ, … Johansson M*. The blood metabolome of incident kidney cancer: A case-
control study nested within the MetKid consortium. PLoS Med. 2021. PMID: 34543281 *Senior author
Studies employing Mendelian randomization
I have long experience in conducting studies where Mendelian randomization has been central in elucidating
modifiable risk factors in cancer etiology, spanning back to my Phd project. At IARC, our research in this area
was accelerated by the availability of large GWA data sets and I have led, or co-led with Dr Brennan, several
MR studies that have improved our understanding of cancer etiology. For instance, MR studies on obesity-
related factors identified elevated body-mass index as an important cause of renal and pancreatic cancer, as
well as highlighted fasting insulin as novel causal risk factor (Carreras-Torres et al. JNCI. 2017, Johansson et
al. PLOS Med. 2019). We further published a study wherein circulating vitamin B12 was evaluated both directly
within the Lung Cancer Cohort Consortium (LC3) and by MR, and observed a notable concordance in the
association between vitamin B12 and histological subtypes of lung cancer (Fanidi et al. Int J Cancer. 2018). In
addition, we conducted a study based on UK Biobank that highlighted that obesity causes a higher uptake and
intensity of tobacco smoking, an observation that may have important implications in tobacco cessation
programs (Carreras-Torres et al. BMJ 2018). Most recently, we carried out a comprehensive analysis using
both MR and direct exposure measurements in large cohort studies to evaluate whether early and mid-life BMI
influences risk of obesity-related cancer (Mariosa et al. JNCI. 2022).
• Carreras-Torres R, Johansson M, Gaborieau V, et al. The Role of Obesity, Type 2 Diabetes, and
Metabolic Factors in Pancreatic Cancer: A Mendelian Randomization Study. J Natl Cancer Inst. 2017.
PMID: 28954281
• Carreras-Torres R*, Johansson M*, Haycock PC, et al. Role of obesity in smoking behaviour: Mendelian
randomisation study in UK Biobank. BMJ. 2018. PMID: 29769355 *Joint first authors
• Fanidi A, Carreras-Torres R, Larose TL, …, Johansson M., Brennan P. Is high vitamin B12 status a
cause of lung cancer? Int J Cancer. 2018. PMID: 30499135
• Mariosa D, Smith-Byrne K, Richardson TG, … Johansson M*. Body Size at Different Ages and Risk of 6
Cancers: A Mendelian Randomization and Prospective Cohort Study. J Natl Cancer Inst. 2022. PMID:
35438160. *Senior author
APPLICATION TO THE ESTONIAN C OMMITTEE ON BIOETH ICS AND HUMAN RESEARCH FOR ETHICAL EVALUATION OF THE RESEARCH PROJECT 1. Name of the study ( in case of a n application in English, the name of the study in Estonian is required in parallel ) Predicting risk of non-communicable disease (PREDICT) Mittenakkushaiguste riski ennustamine (PREDICT) The application is submitted in English, as the researchers are from abroad and do not speak Estonian. 2. The main purpose of the study (up to 450 characters / 0.25 pages) ( in case of a n application in English, the main purpose of the study should be provided in Estonian , too ) Modern medicine has made significant advancements in improving survival for cancer patients, particularly in early detection and novel therapies. Cancer patients in Europe now live nearly six times longer post-diagnosis compared to 40 years ago. Despite these advancements, there is a noticeable lag in extending individuals' "health span"—the time they maintain a high quality of life in good health – without suffering from chronic disease. Up to 40% of cancer cases can be prevented through lifestyle modifications. However, appreciation of how specific behaviours impact individual’s risk of cancer is limited. Existing risk calculators are not able to inform individuals on the total long-term benefits of lifestyle changes across all cancer types. The overarching aim of this project is to establish a unified model for long-term assessment of non-communicable disease ( NCD ) risk that outputs easy-to-understand risk metrics. Amongst NCDs, we will initially consider any cancer, CVD and diabetes, but may subsequently incorporate additional disease endpoints . The main purpose of the study is to: Establish and validate disease-specific risk models based on demographic, behavioural , and anthropometric data. Enrich the standard risk models with risk information from polygenic risk scores and relevant biomarkers. Integrate disease-specific models to predict risk of overall NCD and disease-free survival and evaluate model validity. Main purpose of the study in Estonian: Moodne meditsiin on teinud olulisi edusamme vähipatsientide elulemuse parandamisel, eriti varase tuvastamise ja uute ravivormide osas. Vähiga patsientide eluiga Euroopas on nüüd peaaegu kuus korda pikem pärast diagnoosi saamist võrreldes 40 aasta taguse ajaga. Hoolimata nendest edusammudest on märgatav mahajäämus inimeste tervena elatud aastate pikendamisel – aja, mille jooksul nad säilitavad hea tervise ja elukvaliteedi, ilma krooniliste haigusteta. Kuni 40% vähijuhtudest on võimalik ära hoida elustiili muutuste abil . Siiski on teadlikkus selle kohta, kuidas konkreetsed käitumisviisid mõjutavad inimese vähiriski, piiratud. Praegused riskikalkulaatorid ei võimalda pakkuda inimestele teavet, millist pikaajalist kasu saadakse elustiili muutustest kõigi vähiliikide puhul. Selle projekti kõikehõlmav eesmärk on luua ühtne mudel mittenakkushaiguste (MNH) riskide pikaajaliseks hindamiseks, mis võimaldab tekitada kergesti mõistetavaid riskimõõdikuid . MNH--te hulgas käsitleme esmalt vähki, südame-veresoonkonna haigusi ja diabeeti, kuid hiljem võime lisada täiendavaid haiguste lõpp-punkte. Uuringu põhieesmärk on: Luua ja valideerida haigusspetsiifilised riskimudelid, mis põhinevad demograafilistel, käitumuslikel ja antropomeetrilistel andmetel. Rikastada standardseid riskimudeleid polügeeniliste riskiskooride ja asjakohaste biomarkerite riskiteabega. Integreerida haigusspetsiifilised mudelid, et prognoosida MNH-te ja haigusvaba elu üldisi riske ning hinnata mudeli kehtivust. 3. Principal investigator(s) and their contact details Given name(s): Mattias Last name: Johansson Position: Scientist Institution: International Agency for Research on Cancer (IARC/WHO) Phone: +334 72 73 84 85 e-mail:
[email protected] Skype: NA 4. Other researchers involved in the study ( add lines as necessary ) Given name(s): Allison Last name: Domingues Position: Postdoctoral scientist Institution: International Agency for Research on Cancer (IARC/WHO) Given name(s): Karine Last name: Alcala Position: Data manager Institution: International Agency for Research on Cancer (IARC/WHO) Given name(s): Reedik Last name : Mägi Position: Professor in Bioinformatics Institution : Institute of Genomics, University of Tartu Given name(s): Urmo Last name : Võsa Position: Research Fellow of Functional Genomics Institution : Institute of Genomics, University of Tartu 5. Financing of the study Sources of funding IARC/WHO Total cost of the study ( amount ) Access fees for participating cohorts in this project will be covered by IARC internal budget. Financial compensation for the study participants ( yes, no , explanation and amount ) Not applicable Insurance provided for the study participants ( yes, no, name of the insur ance company and the certificate of insurance (COI)) Not applicable 6. Study period ( the beginning and end dates (MM/YYYY ) ) June 202 5 to May 2029 7. Information about previous or parallel evaluation of the same study project (incl in other countries) NA The study project have received approval from the Scientific Advisory Committee of the Estonian Biobank on 21.11.2024 8. Brief overview of previous studies on the same topic ( up to 900 characters / 0.5 pages) We conducted a preliminary analysis across 16 cancer sites in UK Biobank to quantify the added predictive value of integrating cancer-specific PRS with family history and modifiable risk factors for 16 cancers. We showed that incorporating PRS measurably improves prediction accuracy for most cancers, but the magnitude of this improvement varies substantially. Kachuri L et a. Nat Comm. 2020 Our collaborator (N Chatterjee) developed and validated sex-specific pan-cancer risk scores (PCRSs), defined by the combination of common risk factors and PRSs, to predict the absolute risk of developing at least one of the many common cancer types. This study demonstrated the potential of increasing the risk-benefit balance of rapidly emerging non-invasive multicancer early detection (MCED) liquid biopsy tests through risk stratification. Kim et al. NPJ Prec. Oncol. 2023 Chatterjee et al. further developed a method to establish integrative risk models using incompletely measured risk indicators through Heterogeneous Transfer Learning via GMM. BioRxiv 2023 ; 9. Rationale for the planned study and research questions and / or hypotheses (up to 1800 characters, 1 page) 2027582 1811461 Figure 1. Health trajectories following preventive actions 0 0 Figure 1. Health trajectories following preventive actions Advancements in modern medicine over the past few decades, including in early detection strategies and novel therapies, have significantly improved cancer survival. Cancer patients in Europe live nearly six times longer after their cancer diagnosis today than 40 years ago,[1] and about half of cancer patients survive for 10 years or more after their diagnosis.[2] While these advancements have increased the chronological lifespan of individuals suffering from cancer, they have not necessarily improved their ‘health span’ - how long they remain healthy and free from disease-related morbidity with high quality of life. It is estimated that up to 40% of incident cancers are preventable through health promoting behaviours such as adjustments in diet and physical activity, reduced alcohol consumption, and smoking cessation.[ 3 ] However, though most people have a sense of what a “healthy” lifestyle entails, they may not fully grasp the extent to which lifestyle affects their risk of cancer (Figure 1). Risk calculators, such as the Gail Model for breast cancer, exist to help individuals estimate their cancer risk. However, these calculators focus on individual cancers, while many risk factors such as obesity, smoking and alcohol consumption causally influence the development of several types of cancer, and typically focus on short term risk estimates (5 to 10 years). Without considering the effect of lifestyle factors over multiple diseases and cancers or for a longer time horizon, the total benefit of a shift in lifestyle cannot be assessed using existing tools alone. In addition, existing risk calculators may also fail to present risk information in a way that is understandable to individuals from the public, limiting their ability to influence health behaviours. Increasing health promoting behaviour (e.g., exercising) and reducing modifiable causes (e.g., high blood pressure and obesity), have the potential to decrease the incidence of cancer by up to 50%, [4] and CVDs by 70%. [5] However, such numbers have limited relevance to motivate individuals to reduce their risk of future NCDs. We can readily describe the influence of various risk factors on the short-term risk of developing specific NCDs with the use of disease-specific risk prediction models. Such models are currently used for specific clinical purposes, for instance in (late) prevention of atherosclerotic CVD (ASCVD), and to a lesser extent, risk-based secondary prevention of cancer (i.e., screening) [ 6,7] . However, there are several key challenges to personalised NCD prevention that will be addressed by the CRESCENT initiative, including: (1) Current risk prediction models focus on individual diseases. For instance, a diabetes risk model ignores that the disease shares many risk factors with CVDs and several cancers. This means that the overall benefit of avoiding risk factors, such as obesity or smoking, that causally contribute to multiple diseases is ignored. There is no comprehensive risk prediction model that quantifies the total risk of NCDs and the potential to reduce this risk [8] . (2) Current models provide accurate risk estimates over a relatively short time-period, typically over 5 to 10 years. Such estimates are used by healthcare professionals to guide early detection-based interventions in individuals who are already at high risk for specific diseases. However, such short-term risk estimates do not inform on the long-term benefits of risk reduction through preventive actions earlier in life. (3) Personalised prevention lacks evidence-based risk communication and contextualised guidance. Successful personalised prevention requires carefully developed health counselling and contextualised interventions. Current strategies do not consider individual values or risk literacy, nor social determinants of health and disease, both of which are crucial to determine personal health competencies and self-efficacy. (4) Prospective intervention studies evaluating the effectiveness of implementing risk tools for disease prevention and risk reduction are lacking. Due partially to the lack of a comprehensive prediction model for NCD risk, very few initiatives have translated risk prediction tools into disease prevention or evaluate d their efficacy in prospective interventional studies that measure disease-related endpoints. (5) Policy makers lack an evidence-based tool that assesses population impact from different actions to allow prioritisation of prevention strategies. This is partially owing to the lack of comprehensive tools that can estimate health consequences following implementation of different disease prevention policies. We hypothesise that a ddressing these challenges will require a comprehensive personalised prevention framework, including i) robust prediction models that can accurately estimate overall long-term NCD risk ii) using understandable risk figures to iii) support evidence-based and effective health counselling that can be iv) implemented and adapted to the individual, community, and country to prevent NCDs. Establishing such a risk prediction framework will improve our ability to quantify the impact of disease preventive actions across common NCDs and may also be useful to identify individuals who may benefit from emerging early-detection modalities, such as multi-cancer early detection (MCED). References: R. Berman et al. , “Supportive Care: An Indispensable Component of Modern Oncology,” Clin Oncol , vol. 32, no. 11, pp. 781–788, Nov. 2020, doi : 10.1016/J.CLON.2020.07.020. H. Strongman et al. , “Medium and long-term risks of specific cardiovascular diseases in survivors of 20 adult cancers: a population-based cohort study using multiple linked UK electronic health records databases,” Lancet , vol. 394, no. 10203, p. 1041, Sep. 2019, doi : 10.1016/S0140-6736(19)31674-5. I. Soerjomataram et al. , “Cancers related to lifestyle and environmental factors in France in 2015,” Eur J Cancer , vol. 105, pp. 103–113, Dec. 2018, doi : 10.1016/J.EJCA.2018.09.009. The World Health Organisation. Cancer Prevention. Available at: https://www.who.int/activities/preventing-cancer Yusuf S, Joseph P, Rangarajan S, et al . Modifiable risk factors, cardiovascular disease, and mortality in 155 722 individuals from 21 high-income, middle-income, and low-income countries (PURE): a prospective cohort study. Lancet. 2020 Mar 7;395(10226):795-808. PMID: 31492503 . Tokgozoglu L, Torp-Pedersen C. Redefining cardiovascular risk prediction: is the crystal ball clearer now? Eur Heart J. 2021 Jul 1;42(25):2468-2471. PMID: 34120165. Tammemägi MC. Application of risk prediction models to lung cancer screening: a review. J Thorac Imaging. 2015 Mar;30(2):88-100. PMID: 25692785. Scarborough P, Harrington RA, Mizdrak A, Zhou LM, Doherty A. The Preventable Risk Integrated ModEl and Its Use to Estimate the Health Impact of Public Health Policy Scenarios. Scientifica (Cairo). 2014;2014:748750 . doi : 10.1155/2014/748750. Epub 2014 Sep 25. PMID: 25328757 . 10. Research methodology (up to 1800 characters, 1 page) Data to be obtained: To allow time-to-event modelling of any cancer, CVD and diabetes, we request access to data on all cohort participants (no restrictions) to access d emographic, behavioural, anthropometric, and clinical data . Data will be harmonised among all participants from the Estonian Biobank and other cohorts (EPIC, UKB, HUNT) Design and analysis plan: We will focus on three disease groups that are jointly responsible for more than 60% of the total disease burden in Europe – CVD (leading cause of death, with 37% of all deaths in 2017), cancer (26% of all deaths), and type-2 diabetes (2%). This focus is justified because these NCDs share many modifiable causes that can be intervened upon and measured. Developing the framework for building the NCD prediction model will allow for incorporation of additional NCDs in the future, such as dementia. Furthermore, cancer is heterogenous and the extent to which it can be prevented depends on the site of origin. For instance, lung cancer is highly preventable whereas prostate cancer is not. We will still incorporate prostate cancer and other ‘less preventable’ cancers in the NCD risk model as they represent a significant portion of the cancer burden and may benefit from risk-informed secondary prevention (i.e., early detection). The modelling process will involve i) fitting prediction models for specific diseases (individual cancer types, individual CVDs, and diabetes), ii) enriching the models by incorporating biomarkers information (genetics, proteomics, blood biomarkers, etc.), and iii) integrating these models to predict risk of overall NCD, and disease-free survival. i) For many NCDs there are already well-validated models that can be used to predict disease-specific risk. However, the methods for establishing these models vary between diseases, and whereas that does not invalidate the models, it will be important to ensure that each disease is modelled using a common Cox-based framework to allow estimating the total disease risk. We will therefore refit/recalibrate each disease model using Cox-regression as needed. The diseases include the following ICD codes: all malignant neoplasms excluding non-melanoma skin cancer for prevalent and incident cases (ICD-10 C00-C97 excluding C44); Type 2 diabetes mellitus for prevalent and incident cases (ICD-10: E11); Acute myocardial infarction, subsequent myocardial infarction, complications following acute myocardial infarction, other acute ischaemic heart diseases, chronic ischaemic heart disease, cerebral infarction, stroke not specified as haemorrhage or infarction (I21, I22, I23, I24, I25, I63, I64). ii) We will incorporate additional risk indicators using a data integration and transfer learning methodology developed by our collaborator Nilanjan Chatterjee that allows building models in the setting of partially observed data. The statistical framework entitled Heterogeneous Transfer Learning via GMM has recently been used to establish risk models for multiple diseases in UK Biobank using the Olink proteomics data available on 50,000 research participants, whilst making use of the full cohort database of 500,000 individuals with standard risk indicators. We may also consider incorporating single important, but unmeasured, risk indicators using the Individualised Coherent Absolute Risk Estimator ( iCARE ) method (Chatterjee), which predicts risk by integrating multiple sources of data (risk factor associations with disease, population risk factor distributions, and population rates of disease and mortality). iii) Finally, we will compute cumulative incident functions of each NCD and overall NCD risk by combining the cause-specific hazards estimated in the disease-specific models. Initially, this will involve using simple disease-specific models with key risk indicators (e.g. age, sex, BMI, tobacco and alcohol use) to ensure that the overall risk estimates are well calibrated in the presence of strong competing risks before introducing more complex models. All risk models will be externally validated with respect to their calibration and discrimination. 11. S tudy sample and description of recruitment method. Information and consent forms, questionnaires and tests should be submitted as annexes to the application. Sample size , inclusion of control groups Total number of participants without exclusion in the Estonian B iobank (around 21 2,000), age >18 years old. Who recruits and how/where/ by whom is informed consent obtained? (if applicable) No re-contact or interaction with participants is required. Secondary analysis on available data will be performed. How and from whom are the subjects selected (sampling frame) ? What are inclusion or exclusion criteria of subjects? No re-contact or interaction with participants is required. Secondary analysis on available data will be performed. Type of interventions (physical, mental or data, including special categor ies of personal data) NA Burden on the subject (methods of contact, number of visits, type and number of procedures , repetition of invitations, etc.) NA 12. Issuing of tissue samples to third parties (RNA, DNA, plasma etc ) The number of gene donors whose tissue samples will be issued and the types of tissue samples to be issued NA The amount of tissue sample to be issued per one gene donor NA The entity to whom tissue samples will be issued ( country, institution , a d dress)? NA What will be done with the residue samples ( will the residue samples be destroyed or sent back to Gene Bank )? NA 1 3 . Analysis of the ethical aspects of the study (3600 characters, up to 2 pages). All research involving human s ubjects must be carried out in compliance with ethical requirements, in particular the principles of respect for autonomy, charity and the prevention of harm, and justice. ( https://www.coe.int/en/web/bioethics/guide-for-research-ethics-committees-members ). The researchers involved in data analysis are employed by the International Agency for Research on Cancer (IARC/WHO). IARC/WHO is committed to upholding the highest ethical standards in all research activities, including those involving human biological samples, data protection, and privacy concerns. While IARC/WHO is not bound by national or regional data protection laws due to its international status, it ensures that human data is handled in line with internationally recognised data protection frameworks. This includes compliance with the Personal Data Protection and Privacy Principles for UN System Organisations (UN-HCLM 2018) and the IARC Data Protection Policy . In this study, researchers will access data remotely from the Estonian Biobank . The following ethical principles have been specifically considered in the planning of this study: Respect for Individual Autonomy Respect for autonomy is a core ethical tenet of this study. All participants in the Estonian Biobank have provided informed consent to participate in research, including the use of their health data and genetic information. The data will be accessed in pseudonymised form, ensuring participant identities are protected. No re-identification or re-encoding of the data will be permitted. Researchers will access the data remotely, and it will be used exclusively for modelling and predicting risk of cancer and other non-communicable diseases (NCDs), with an emphasis on the long-term effects of lifestyle changes. Non-Maleficence (Do No Harm) The principle of non-maleficence underpins the entire project design. The study aims to develop and validate risk models for cancer, cardiovascular disease (CVD), and diabetes using demographic, behavioural, and anthropometric data. Advanced methodologies, including polygenic risk scores and biomarker integration, will be employed to ensure scientific rigour and accuracy. The study is designed to minimise risk and maximise benefit by providing insights that can inform personalised prevention strategies. No interventions will be carried out directly on participants, and all data will remain securely managed, significantly reducing the potential for harm. Justice The Estonian Biobank represents a population-based cohort encompassing a diverse cross-section of Estonian society, including variation in age, sex, and socioeconomic background. This diversity enhances the generalisability of the findings. By developing integrated risk models that estimate the likelihood of various NCDs and predict disease-free survival, the research aims to support public health by improving risk communication and prevention strategies across the population. Ultimately, the study seeks to deliver equitable health benefits, contributing to the reduction of NCD burden in a fair and inclusive manner. 1 3 a Human s ubjects Assistance questions No Yes Are people the object of research? Secondary analysis o f pseudonymised individual level data will be performed . All of the participants have joined the Estonian Biobank voluntarily and given informed consent for data analysis for the purpose of scientific research. Nobody will be discriminated against upon joining or for the fact of being a participant. The participant can withdraw consent at any time. Gene donors can prohibit the use of other databases containing records of the donor. Are the study participants vulnerable individuals or groups ? X Does the study include persons who cannot themselves give informed consent to part icipate in the research (incl. p ersons with limited active legal capacity)? X Are the study participants children/ minors? X Are the study participants patients ? X Does the research involve collection of biological samples? Are human biological samples intende d for export to a third country ( https://www.aki.ee/et/teenused-poordumisvormid/andmete-edastamine-valisriiki ) ( https://www.aki.ee/en/guidelines/transfer-personal-data-foreign-country-0 ) or import them from another country to Estonia? X 1 3 b Personal data and datasets No Yes Are personal data collected or analyzed in the study , including special categories of personal data? Pseudonymised previously collected individual level data on study participants will be analyzed, including risk factor information, g enomic data , and disease information. Does the research involve systematic monitoring of an individual, the collection of his or her data profile, or a large-scale processing of data of special categories and /or sensitive data, or the use of (intrusive) data processing techniques in a covert way (eg survival surveys, monitoring, surveillance, audio and video recording, geolocation, etc.) or any data processing that may harm the rights and freedoms of the data subject? X Is there a plan to analyze previously collected personal data? Pseudonymised previously collected individual level data on study participants will be analyzed, including risk factor information, genomic data, and disease information. Is there a plan to analyze publicly available data? X Is there an intention to transfer personal data or provide access to personal data to third countries ( https://www.aki.ee/et/teenused-poordumisvormid/andmete-edastamine-valisriiki )? ( https://www.aki.ee/en/guidelines/transfer-personal-data-foreign-country-0 ) X Will personal data be destroyed / anonymised at the end of the research? X 1 3 c Other ethical issues Can conducting research involve ethical risks not described above? X 14. Complete in case the research is based on data from a database and/or register Name of database and/or register The Estonian Biobank database Purpose of the processing of personal data The purpose of processing is to develop and validate risk models for non-chronical diseases. List of variables and period for which data are collected ( in annex if necessary) We will analyze data in the following categories G enealogical data (family history of medical conditions spanning four generations , if available ) including: Cancers (breast, brain, bowel, gastrointestinal, liver, lung, lymphoma, myeloma, prostate, and stomach/gastric) Cardiovascular disease Diabetes High blood pressure High cholesterol E ducational and occupational history : Highest degree completed Occupational history Occupational exposures (paint, asbestos, pesticides) L ifestyle data including: P hysical activity S moking information (status, intensity, duration, quit years in former smokers) A lcohol information (status, intensity, duration) D ietary information (processed meat consumption, fruit and vegetable intake, salt intake) Other dietary information (FFQ etc) residence data at county level (to assess indirectly pollution exposure ) Women’s health information including: Age at menarche Menopausal status Use of menopausal hormonal therapy Contraception use (type and duration) Parity Age at first live birth Other health measurements: Weight, height HbA1c PSA Blood pressure Cholesterol levels Pulse rate Triglycerides Glucose Uric acid C-reactive protein Genetic data Clinical information on incident and prevalent disease including: Allergic conditions Diabetes High cholesterol Hypertension Hepatitis C/B Non-viral liver disease GERD Chronic kidney disease COPD Cardiovascular disease Cancers 15 . Description of personal data protection measures, including data storage, security and erasure, including date of erasure of data and / or code key (up to 1800 characters, 1 page). Describe and justify the storage of data collected for the study and the deadline for storage . All data used in this study will be accessed remotely through the SAPU environment in a pseudonymised format. No identifiable personal data will be transferred or downloaded. Researchers from IARC/WHO will access the data through a secure, encrypted platform provided by the University of Tartu , which includes strict access controls and user authentication. IARC/WHO is committed to safeguarding personal data in accordance with international standards, including the UN Personal Data Protection and Privacy Principles (2018) and the IARC Data Protection Policy. Although not subject to national legislation, IARC applies rigorous internal protocols to ensure data privacy, confidentiality, and security. Data will be used solely for the purpose of this study and will not be depseudonymised . No copies of the data will be stored outside the secure platform. Upon completion of the study, access rights will be revoked in line with the Estonian Biobank’s policies. The code key linking pseudonyms to personal identities is held exclusively by the Estonian Biobank and is not accessible to IARC researchers at any stage. At the end of the project, a data audit is conducted to decide which data to keep or archive and which to delete. After that, the SAPU environment is deleted. Deadline: May 2029 Describe the process and means of pseudonymisation of personal data. Pseudonymisation is performed by Estonian Biobank. Is there a plan to de-pseudonymise gene donors’ data? Please specify the number of gene donors whose data will be de-pseudonymised. Please explain the reason for de-pseudonymisation. Is there a plan to transport personal data ? Please describe how data protection is ensured. Personal data will be analysed in SAPU . Describe how the data are protected against unauthorized or unlawful processing. The Sensitive D ata A nalysis P latform or SAPU is an environment provided by the University of Tartu High-Performance Computing Centre, where analysts and programmers can work on sensitive data https://docs.hpc.ut.ee/public/services/SAPU/ The environment reduces the risk of possible unauthorized copy, transfer, or retrieval of sensitive data from the machines, providing a higher class of security than that of a standard high-performance cluster. SAPU is an isolated environment where: The machine has no direct access to the internet. Complete network isolation based on firewall rules. Access to the machine is possible only through a virtual desktop environment. Analysts can move files using object storage, which saves anything moved. Moving files out requires approval from the data owners' side. The monitoring layer and the server record all actions taken. SAPU supports most research tools available to scientist . I confirm that all researchers are aware of the ethical and personal data protection requirements of the project. Signature of the principal investigator Date of application 19 .05. 2025 3753485 5930900 0 0 3467735 5788660 0 0 3467735 5788660 0 0 EBIN ID of the application ( fills by the assessor ) List of additional documents: 1. CV of the principal investigator
Saatja: "Shaymaa Alwaheidi" <
[email protected]>
Saaja: "Info - SOM" <
[email protected]>
Teema: EBIN Application/ PREDICT project (IARC/WHO)
Kuupäev: 2025-05-19 11:34
Tähelepanu! Tegemist on välisvõrgust saabunud kirjaga.
Tundmatu saatja korral palume linke ja faile mitte avada.
Dear Office of the Ethics Committee of the Estonian Ministry of Social
Affairs (EBIN),
We would like to submit our ethics application for the project titled
Predicting Risk of Non-Communicable Disease (PREDICT) / Mittenakkushaiguste
Riski Ennustamine (PREDICT).
As the research team is based abroad and does not speak Estonian, the
application has been submitted in English. I have attached the approval
decision from the Scientific Advisory Committee of the Estonian Biobank,
along with the CV of the principal investigator.
If you need anything else, please feel free to reach out.
Best wishes,
Shaymaa, Allison, and Mattias
NOTICE OF CONFIDENTIALITY AND/OR LEGAL PRIVILEGE
This e-mail message contains proprietary information of the International
Agency for Research on Cancer, an agency of the World Health Organization,
that is strictly confidential and/or that is legally privileged. It is
intended solely for the use of officials of the International Agency for
Research on Cancer/ World Health Organization, and/or the named recipient(s)
of the message. ANY UNAUTHORIZED REVIEW, DISCLOSURE, COPYING, DISTRIBUTION,
RELIANCE ON THE CONTENTS, OR OTHER USE OF THE INFORMATION IN THIS MESSAGE IS
STRICTLY PROHIBITED. If you have received this message in error, please
notify the sender immediately and permanently delete this message, any
attachments, and any copies of it from your e-mail system.
DECISION No 1 OF THE SCIENTIFIC ADVISORY COMMITTEE OF THE ESTONIAN BIOBANK
21.11.2024
The Scientific Advisory Committee of the Estonian Biobank
Chair: Elin Org
Members:
1. Lili Milani (Head of the Estonian Biobank)
2. Priit Palta (Head of the Development Centre of the Estonian Biobank)
3. Elin Org (Head of the Research Centre of the Estonian Biobank)
4. Kristjan Metsalu (Head of the IT Department of the Estonian Biobank)
5. Reedik Mägi (Professor of Bioinformatics)
6. Neeme Tõnisson (Professor of Medical Genetics)
7. Helene Alavere (Project Manager and Head of the Data Collection Department, Institute
of Genomics, UT)
8. Steven Smit (Head of the Biobank Laboratory of the Estonian Biobank)
9. Urmo Võsa (Research Fellow in Functional Genomics)
discussed at its meeting on 21.11.2024 the proposal of Mattias Johansson, scientist of the
International Agency for Research on Cancer (IARC/WHO), to conduct the study " Predicting risk
of non-communicable disease (PREDICT)".
DECISION: The application is approved on the condition that data is accessed via secure
SAPU servers.
We recommend contacting researchers involved in ongoing scientific projects using similar data
to facilitate the fast utilisation of data and the smooth conduct of your study. For CVD and blood
biomarkers, Urmo Võsa (
[email protected]), and for T2D and cancer, Reedik Mägi
(
[email protected]).
Elin Org
Chair of the Scientific Advisory Committee
/Digitally signed/