Publications

17.

Wicker, Jörg; Tyukin, Andrey; Kramer, Stefan

A Nonlinear Label Compression and Transformation Method for Multi-Label Classification using Autoencoders Proceedings Article

In: Bailey, James; Khan, Latifur; Washio, Takashi; Dobbie, Gill; Huang, Zhexue Joshua; Wang, Ruili (Ed.): The 20th Pacific Asia Conference on Knowledge Discovery and Data Mining (PAKDD), pp. 328-340, Springer International Publishing, Switzerland, 2016, ISBN: 978-3-319-31753-3.

@inproceedings{wicker2016nonlinear,

title = {A Nonlinear Label Compression and Transformation Method for Multi-Label Classification using Autoencoders},

author = {J\"{o}rg Wicker and Andrey Tyukin and Stefan Kramer},

editor = {James Bailey and Latifur Khan and Takashi Washio and Gill Dobbie and Zhexue Joshua Huang and Ruili Wang},

url = {http://dx.doi.org/10.1007/978-3-319-31753-3_27},

doi = {10.1007/978-3-319-31753-3_27},

isbn = {978-3-319-31753-3},

year  = {2016},

date = {2016-04-16},

booktitle = {The 20th Pacific Asia Conference on Knowledge Discovery and Data Mining (PAKDD)},

volume = {9651},

pages = {328-340},

publisher = {Springer International Publishing},

address = {Switzerland},

series = {Lecture Notes in Computer Science},

abstract = {Multi-label classification targets the prediction of multiple interdependent and non-exclusive binary target variables. Transformation-based algorithms transform the data set such that regular single-label algorithms can be applied to the problem. A special type of transformation-based classifiers are label compression methods, that compress the labels and then mostly use single label classifiers to predict the compressed labels. So far, there are no compression-based algorithms follow a problem transformation approach and address non-linear dependencies in the labels. In this paper, we propose a new algorithm, called Maniac (Multi-lAbel classificatioN usIng AutoenCoders), which extracts the non-linear dependencies by compressing the labels using autoencoders. We adapt the training process of autoencoders in a way to make them more suitable for a parameter optimization in the context of this algorithm. The method is evaluated on eight standard multi-label data sets. Experiments show that despite not producing a good ranking, Maniac generates a particularly good bipartition of the labels into positives and negatives. This is caused by rather strong predictions with either really high or low probability. Additionally, the algorithm seems to perform better given more labels and a higher label cardinality in the data set.},

keywords = {autoencoders, label compression, machine learning, multi-label classification},

pubstate = {published},

tppubtype = {inproceedings}

}

Close

16.

Wicker, Jörg; Lorsbach, Tim; Gütlein, Martin; Schmid, Emanuel; Latino, Diogo; Kramer, Stefan; Fenner, Kathrin

enviPath – The Environmental Contaminant Biotransformation Pathway Resource Journal Article

In: Nucleic Acid Research, vol. 44, no. D1, pp. D502-D508, 2016.

15.

Raza, Atif; Wicker, Jörg; Kramer, Stefan

Trading Off Accuracy for Efficiency by Randomized Greedy Warping Proceedings Article

In: Proceedings of the 31st Annual ACM Symposium on Applied Computing, pp. 883-890, ACM, New York, NY, USA, 2016, ISBN: 978-1-4503-3739-7.

14.

Williams, Jonathan; Stönner, Christof; Wicker, Jörg; Krauter, Nicolas; Derstorff, Bettina; Bourtsoukidis, Efstratios; Klüpfel, Thomas; Kramer, Stefan

Cinema audiences reproducibly vary the chemical composition of air during films, by broadcasting scene specific emissions on breath Journal Article

In: Scientific Reports, vol. 6, 2016.

13.

Wicker, Jörg; Krauter, Nicolas; Derstorff, Bettina; Stönner, Christof; Bourtsoukidis, Efstratios; Klüpfel, Thomas; Williams, Jonathan; Kramer, Stefan

Cinema Data Mining: The Smell of Fear Proceedings Article

In: Proceedings of the 21st ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 1235-1304, ACM ACM, New York, NY, USA, 2015, ISBN: 978-1-4503-3664-2.

@inproceedings{wicker2015cinema,

title = {Cinema Data Mining: The Smell of Fear},

author = {J\"{o}rg Wicker and Nicolas Krauter and Bettina Derstorff and Christof St\"{o}nner and Efstratios Bourtsoukidis and Thomas Kl\"{u}pfel and Jonathan Williams and Stefan Kramer},

url = {https://wicker.nz/nwp-acm/authorize.php?id=N10031 

http://doi.acm.org/10.1145/2783258.2783404},

doi = {10.1145/2783258.2783404},

isbn = {978-1-4503-3664-2},

year  = {2015},

date = {2015-01-01},

booktitle = {Proceedings of the 21st ACM SIGKDD International Conference on Knowledge Discovery and Data Mining},

pages = {1235-1304},

publisher = {ACM},

address = {New York, NY, USA},

organization = {ACM},

series = {KDD '15},

abstract = {While the physiological response of humans to emotional events or stimuli is well-investigated for many modalities (like EEG, skin resistance, ...), surprisingly little is known about the exhalation of so-called Volatile Organic Compounds (VOCs) at quite low concentrations in response to such stimuli. VOCs are molecules of relatively small mass that quickly evaporate or sublimate and can be detected in the air that surrounds us. The paper introduces a new field of application for data mining, where trace gas responses of people reacting on-line to films shown in cinemas (or movie theaters) are related to the semantic content of the films themselves. To do so, we measured the VOCs from a movie theatre over a whole month in intervals of thirty seconds, and annotated the screened films by a controlled vocabulary compiled from multiple sources. To gain a better understanding of the data and to reveal unknown relationships, we have built prediction models for so-called forward prediction (the prediction of future VOCs from the past), backward prediction (the prediction of past scene labels from future VOCs) and for some forms of abductive reasoning and Granger causality. Experimental results show that some VOCs and some labels can be predicted with relatively low error, and that hints for causality with low p-values can be detected in the data.},

keywords = {atmospheric chemistry, breath analysis, causality, cheminformatics, cinema data mining, data mining, emotional response analysis, movie analysis, smell of fear, sof, time series},

pubstate = {published},

tppubtype = {inproceedings}

}

Close

12.

Tyukin, Andrey; Kramer, Stefan; Wicker, Jörg

Scavenger – A Framework for the Efficient Evaluation of Dynamic and Modular Algorithms Proceedings Article

In: Bifet, Albert; May, Michael; Zadrozny, Bianca; Gavalda, Ricard; Pedreschi, Dino; Cardoso, Jaime; Spiliopoulou, Myra (Ed.): Machine Learning and Knowledge Discovery in Databases, pp. 325-328, Springer International Publishing, 2015, ISBN: 978-3-319-23460-1.

11.

Tyukin, Andrey; Kramer, Stefan; Wicker, Jörg

BMaD — A Boolean Matrix Decomposition Framework Proceedings Article

In: Calders, Toon; Esposito, Floriana; Hüllermeier, Eyke; Meo, Rosa (Ed.): Machine Learning and Knowledge Discovery in Databases, pp. 481-484, Springer Berlin Heidelberg, 2014, ISBN: 978-3-662-44844-1.

10.

Wicker, Jörg

Large Classifier Systems in Bio- and Cheminformatics PhD Thesis

Technische Universität München, 2013.

Abstract | Links | BibTeX | Tags: biodegradation, bioinformatics, cheminformatics, computational sustainability, data mining, enviPath, machine learning, multi-label classification, multi-relational learning, toxicity

@phdthesis{wicker2013large,

title = {Large Classifier Systems in Bio- and Cheminformatics},

author = {J\"{o}rg Wicker},

url = {http://mediatum.ub.tum.de/node?id=1165858},

year  = {2013},

date = {2013-01-01},

school = {Technische Universit\"{a}t M\"{u}nchen},

abstract = {Large classifier systems are machine learning algorithms that use multiple 

classifiers to improve the prediction of target values in advanced 

classification tasks. Although learning problems in bio- and 

cheminformatics commonly provide data in schemes suitable for large 

classifier systems, they are rarely used in these domains. This thesis 

introduces two new classifiers incorporating systems of classifiers 

using Boolean matrix decomposition to handle data in a schema that 

often occurs in bio- and cheminformatics. 

 

The first approach, called MLC-BMaD (multi-label classification using 

Boolean matrix decomposition), uses Boolean matrix decomposition to 

decompose the labels in a multi-label classification task. The 

decomposed matrices are a compact representation of the information 

in the labels (first matrix) and the dependencies among the labels 

(second matrix). The first matrix is used in a further multi-label 

classification while the second matrix is used to generate the final 

matrix from the predicted values of the first matrix. 

MLC-BMaD was evaluated on six standard multi-label data sets, the 

experiments showed that MLC-BMaD can perform particularly well on data 

sets with a high number of labels and a small number of instances and 

can outperform standard multi-label algorithms. 

Subsequently, MLC-BMaD is extended to a special case of 

multi-relational learning, by considering the labels not as simple 

labels, but instances. The algorithm, called ClassFact 

(Classification factorization), uses both matrices in a multi-label 

classification. Each label represents a mapping between two 

instances. 

Experiments on three data sets from the domain of bioinformatics show 

that ClassFact can outperform the baseline method, which merges the 

relations into one, on hard classification tasks. 

 

Furthermore, large classifier systems are used on two cheminformatics 

data sets, the first one is used to predict the environmental fate of 

chemicals by predicting biodegradation pathways. The second is a data 

set from the domain of predictive toxicology. In biodegradation 

pathway prediction, I extend a knowledge-based system and incorporate 

a machine learning approach to predict a probability for 

biotransformation products based on the structure- and knowledge-based 

predictions of products, which are based on transformation rules. The 

use of multi-label classification improves the performance of the 

classifiers and extends the number of transformation rules that can be 

covered. 

For the prediction of toxic effects of chemicals, I applied large 

classifier systems to the ToxCasttexttrademark data set, which maps 

toxic effects to chemicals. As the given toxic effects are not easy to 

predict due to missing information and a skewed class 

distribution, I introduce a filtering step in the multi-label 

classification, which finds labels that are usable in multi-label 

prediction and does not take the others in the 

prediction into account. Experiments show 

that this approach can improve upon the baseline method using binary 

classification, as well as multi-label approaches using no filtering. 

 

The presented results show that large classifier systems can play a 

role in future research challenges, especially in bio- and 

cheminformatics, where data sets frequently consist of more complex 

structures and data can be rather small in terms of the number of 

instances compared to other domains.},

keywords = {biodegradation, bioinformatics, cheminformatics, computational sustainability, data mining, enviPath, machine learning, multi-label classification, multi-relational learning, toxicity},

pubstate = {published},

tppubtype = {phdthesis}

}

Close

Large classifier systems are machine learning algorithms that use multiple
classifiers to improve the prediction of target values in advanced
classification tasks. Although learning problems in bio- and
cheminformatics commonly provide data in schemes suitable for large
classifier systems, they are rarely used in these domains. This thesis
introduces two new classifiers incorporating systems of classifiers
using Boolean matrix decomposition to handle data in a schema that
often occurs in bio- and cheminformatics.

The first approach, called MLC-BMaD (multi-label classification using
Boolean matrix decomposition), uses Boolean matrix decomposition to
decompose the labels in a multi-label classification task. The
decomposed matrices are a compact representation of the information
in the labels (first matrix) and the dependencies among the labels
(second matrix). The first matrix is used in a further multi-label
classification while the second matrix is used to generate the final
matrix from the predicted values of the first matrix.
MLC-BMaD was evaluated on six standard multi-label data sets, the
experiments showed that MLC-BMaD can perform particularly well on data
sets with a high number of labels and a small number of instances and
can outperform standard multi-label algorithms.
Subsequently, MLC-BMaD is extended to a special case of
multi-relational learning, by considering the labels not as simple
labels, but instances. The algorithm, called ClassFact
(Classification factorization), uses both matrices in a multi-label
classification. Each label represents a mapping between two
instances.
Experiments on three data sets from the domain of bioinformatics show
that ClassFact can outperform the baseline method, which merges the
relations into one, on hard classification tasks.

Furthermore, large classifier systems are used on two cheminformatics
data sets, the first one is used to predict the environmental fate of
chemicals by predicting biodegradation pathways. The second is a data
set from the domain of predictive toxicology. In biodegradation
pathway prediction, I extend a knowledge-based system and incorporate
a machine learning approach to predict a probability for
biotransformation products based on the structure- and knowledge-based
predictions of products, which are based on transformation rules. The
use of multi-label classification improves the performance of the
classifiers and extends the number of transformation rules that can be
covered.
For the prediction of toxic effects of chemicals, I applied large
classifier systems to the ToxCasttexttrademark data set, which maps
toxic effects to chemicals. As the given toxic effects are not easy to
predict due to missing information and a skewed class
distribution, I introduce a filtering step in the multi-label
classification, which finds labels that are usable in multi-label
prediction and does not take the others in the
prediction into account. Experiments show
that this approach can improve upon the baseline method using binary
classification, as well as multi-label approaches using no filtering.

The presented results show that large classifier systems can play a
role in future research challenges, especially in bio- and
cheminformatics, where data sets frequently consist of more complex
structures and data can be rather small in terms of the number of
instances compared to other domains.

Close

9.

Wicker, Jörg; Pfahringer, Bernhard; Kramer, Stefan

Multi-label Classification Using Boolean Matrix Decomposition Proceedings Article

In: Proceedings of the 27th Annual ACM Symposium on Applied Computing, pp. 179–186, ACM, 2012, ISBN: 978-1-4503-0857-1.

8.

Hardy, Barry; Douglas, Nicki; Helma, Christoph; Rautenberg, Micha; Jeliazkova, Nina; Jeliazkov, Vedrin; Nikolova, Ivelina; Benigni, Romualdo; Tcheremenskaia, Olga; Kramer, Stefan; Girschick, Tobias; Buchwald, Fabian; Wicker, Jörg; Karwath, Andreas; Gütlein, Martin; Maunz, Andreas; Sarimveis, Haralambos; Melagraki, Georgia; Afantitis, Antreas; Sopasakis, Pantelis; Gallagher, David; Poroikov, Vladimir; Filimonov, Dmitry; Zakharov, Alexey; Lagunin, Alexey; Gloriozova, Tatyana; Novikov, Sergey; Skvortsova, Natalia; Druzhilovsky, Dmitry; Chawla, Sunil; Ghosh, Indira; Ray, Surajit; Patel, Hitesh; Escher, Sylvia

Collaborative development of predictive toxicology applications Journal Article

In: Journal of Cheminformatics, vol. 2, no. 1, pp. 7, 2010, ISSN: 1758-2946.

@article{hardy2010collaborative,

title = {Collaborative development of predictive toxicology applications},

author = {Barry Hardy and Nicki Douglas and Christoph Helma and Micha Rautenberg and Nina Jeliazkova and Vedrin Jeliazkov and Ivelina Nikolova and Romualdo Benigni and Olga Tcheremenskaia and Stefan Kramer and Tobias Girschick and Fabian Buchwald and J\"{o}rg Wicker and Andreas Karwath and Martin G\"{u}tlein and Andreas Maunz and Haralambos Sarimveis and Georgia Melagraki and Antreas Afantitis and Pantelis Sopasakis and David Gallagher and Vladimir Poroikov and Dmitry Filimonov and Alexey Zakharov and Alexey Lagunin and Tatyana Gloriozova and Sergey Novikov and Natalia Skvortsova and Dmitry Druzhilovsky and Sunil Chawla and Indira Ghosh and Surajit Ray and Hitesh Patel and Sylvia Escher},

url = {http://www.jcheminf.com/content/2/1/7},

doi = {10.1186/1758-2946-2-7},

issn = {1758-2946},

year  = {2010},

date = {2010-01-01},

journal = {Journal of Cheminformatics},

volume = {2},

number = {1},

pages = {7},

abstract = {OpenTox provides an interoperable, standards-based Framework for the support of predictive toxicology data management, algorithms, modelling, validation and reporting. It is relevant to satisfying the chemical safety assessment requirements of the REACH legislation as it supports access to experimental data, (Quantitative) Structure-Activity Relationship models, and toxicological information through an integrating platform that adheres to regulatory requirements and OECD validation principles. Initial research defined the essential components of the Framework including the approach to data access, schema and management, use of controlled vocabularies and ontologies, architecture, web service and communications protocols, and selection and integration of algorithms for predictive modelling. OpenTox provides end-user oriented tools to non-computational specialists, risk assessors, and toxicological experts in addition to Application Programming Interfaces (APIs) for developers of new applications. OpenTox actively supports public standards for data representation, interfaces, vocabularies and ontologies, Open Source approaches to core platform components, and community-based collaboration approaches, so as to progress system interoperability goals.The OpenTox Framework includes APIs and services for compounds, datasets, features, algorithms, models, ontologies, tasks, validation, and reporting which may be combined into multiple applications satisfying a variety of different user needs. OpenTox applications are based on a set of distributed, interoperable OpenTox API-compliant REST web services. The OpenTox approach to ontology allows for efficient mapping of complementary data coming from different datasets into a unifying structure having a shared terminology and representation.Two initial OpenTox applications are presented as an illustration of the potential impact of OpenTox for high-quality and consistent structure-activity relationship modelling of REACH-relevant endpoints: ToxPredict which predicts and reports on toxicities for endpoints for an input chemical structure, and ToxCreate which builds and validates a predictive toxicity model based on an input toxicology dataset. Because of the extensible nature of the standardised Framework design, barriers of interoperability between applications and content are removed, as the user may combine data, models and validation from multiple sources in a dependable and time-effective way.},

keywords = {cheminformatics, computational sustainability, data mining, machine learning, REST, toxicity},

pubstate = {published},

tppubtype = {article}

}

Close

OpenTox provides an interoperable, standards-based Framework for the support of predictive toxicology data management, algorithms, modelling, validation and reporting. It is relevant to satisfying the chemical safety assessment requirements of the REACH legislation as it supports access to experimental data, (Quantitative) Structure-Activity Relationship models, and toxicological information through an integrating platform that adheres to regulatory requirements and OECD validation principles. Initial research defined the essential components of the Framework including the approach to data access, schema and management, use of controlled vocabularies and ontologies, architecture, web service and communications protocols, and selection and integration of algorithms for predictive modelling. OpenTox provides end-user oriented tools to non-computational specialists, risk assessors, and toxicological experts in addition to Application Programming Interfaces (APIs) for developers of new applications. OpenTox actively supports public standards for data representation, interfaces, vocabularies and ontologies, Open Source approaches to core platform components, and community-based collaboration approaches, so as to progress system interoperability goals.The OpenTox Framework includes APIs and services for compounds, datasets, features, algorithms, models, ontologies, tasks, validation, and reporting which may be combined into multiple applications satisfying a variety of different user needs. OpenTox applications are based on a set of distributed, interoperable OpenTox API-compliant REST web services. The OpenTox approach to ontology allows for efficient mapping of complementary data coming from different datasets into a unifying structure having a shared terminology and representation.Two initial OpenTox applications are presented as an illustration of the potential impact of OpenTox for high-quality and consistent structure-activity relationship modelling of REACH-relevant endpoints: ToxPredict which predicts and reports on toxicities for endpoints for an input chemical structure, and ToxCreate which builds and validates a predictive toxicity model based on an input toxicology dataset. Because of the extensible nature of the standardised Framework design, barriers of interoperability between applications and content are removed, as the user may combine data, models and validation from multiple sources in a dependable and time-effective way.

Close

7.

Wicker, Jörg; Fenner, Kathrin; Ellis, Lynda; Wackett, Larry; Kramer, Stefan

Predicting biodegradation products and pathways: a hybrid knowledge- and machine learning-based approach Journal Article

In: Bioinformatics, vol. 26, no. 6, pp. 814-821, 2010.

6.

Wicker, Jörg; Richter, Lothar; Kramer, Stefan

SINDBAD and SiQL: Overview, Applications and Future Developments Book Section

In: Džeroski, Sašo; Goethals, Bart; Panov, Panče (Ed.): Inductive Databases and Constraint-Based Data Mining, pp. 289-309, Springer New York, 2010, ISBN: 978-1-4419-7737-3.

5.

Wicker, Jörg; Richter, Lothar; Kessler, Kristina; Kramer, Stefan

SINDBAD and SiQL: An Inductive Database and Query Language in the Relational Model Proceedings Article

In: Daelemans, Walter; Goethals, Bart; Morik, Katharina (Ed.): Machine Learning and Knowledge Discovery in Databases, pp. 690-694, Springer Berlin Heidelberg, 2008, ISBN: 978-3-540-87480-5.

4.

Richter, Lothar; Wicker, Jörg; Kessler, Kristina; Kramer, Stefan

An Inductive Database and Query Language in the Relational Model Proceedings Article

In: Proceedings of the 11th International Conference on Extending Database Technology: Advances in Database Technology, pp. 740–744, ACM, 2008, ISBN: 978-1-59593-926-5.

3.

Wicker, Jörg; Brosdau, Christoph; Richter, Lothar; Kramer, Stefan

SINDBAD SAILS: A Service Architecture for Inductive Learning Schemes Proceedings Article

In: Proceedings of the First Workshop on Third Generation Data Mining: Towards Service-Oriented Knowledge Discovery, 2008.

Abstract | Links | BibTeX | Tags: data mining, inductive databases, machine learning, query languages

2.

Wicker, Jörg; Fenner, Kathrin; Ellis, Lynda; Wackett, Larry; Kramer, Stefan

Machine Learning and Data Mining Approaches to Biodegradation Pathway Prediction Proceedings Article

In: Bridewell, Will; Calders, Toon; Medeiros, Ana Karla; Kramer, Stefan; Pechenizkiy, Mykola; Todorovski, Ljupco (Ed.): Proceedings of the Second International Workshop on the Induction of Process Models at ECML PKDD 2008, 2008.

Links | BibTeX | Tags: biodegradation, cheminformatics, computational sustainability, enviPath, machine learning, metabolic pathways

1.

Kramer, Stefan; Aufschild, Volker; Hapfelmeier, Andreas; Jarasch, Alexander; Kessler, Kristina; Reckow, Stefan; Wicker, Jörg; Richter, Lothar

Inductive Databases in the Relational Model: The Data as the Bridge Proceedings Article

In: Bonchi, Francesco; Boulicaut, Jean-François (Ed.): Knowledge Discovery in Inductive Databases, pp. 124-138, Springer Berlin Heidelberg, 2006, ISBN: 978-3-540-33292-3.