We meet on Wednesdays at 1pm, in the 10th floor conference room of the Statistics Department, 1255 Amsterdam Ave, New York, NY.
Friday, August 29, 2014
Monday, July 7, 2014
Daniel Soudry: July 9th
We will discuss the following book chapter:
F. Bach, R. Jenatton, J. Mairal and G. Obozinski. Convex optimization with sparsity-inducing norms. In S. Sra, S. Nowozin, S. J. Wright., editors, Optimization for Machine Learning, MIT Press, 2011.
http://www.di.ens.fr/~fbach/opt_book.pdf
F. Bach, R. Jenatton, J. Mairal and G. Obozinski. Convex optimization with sparsity-inducing norms. In S. Sra, S. Nowozin, S. J. Wright., editors, Optimization for Machine Learning, MIT Press, 2011.
http://www.di.ens.fr/~fbach/opt_book.pdf
Monday, June 30, 2014
Vamsi Krishna Potluru: July 2nd
Efficient Sparse NMF for fMRI data analysis
Nonnegative matrix factorization (NMF) has become a ubiquitous tool for data analysis. An important variant is the sparse NMF problem which arises when we explicitly require the learnt features to be sparse. A natural measure of sparsity is the L0 norm, however its optimization is NP-hard. Mixed norms, such as L1/L2 measure, have been shown to model sparsity robustly, based on intuitive attributes that such measures need to satisfy. This is in contrast to computationally cheaper alternatives such as the plain L1 norm. However, present algorithms designed for optimizing the mixed norm L1/L2 are slow and other formulations for sparse NMF have been
proposed such as those based on L1 and L0 norms. Our proposed algorithm allows us to solve the mixed norm sparsity constraints while not sacri ficing computation time. We present experimental evidence on real-world datasets that
shows our new algorithm performs an order of magnitude faster compared
to the current state-of-the-art solvers optimizing the mixed norm and is suitable for large-scale datasets [1]. Also,
recently, its computational efficiency has been exploited for evaluating
the sparse NMF model for fMRI analysis [2]. And the authors show that
the sparse NMF model is competitive with other state-of-the-art matrix
factorization methods such as ICA, sparse PCA and even restricted
Boltzmann machines.
Links:
Home: http://www.vamsi.guru
Thursday, June 19, 2014
Friday, May 30, 2014
Josh Merel: June 4th
Josh will present about linear matrix inequalities and their relevance for control theory problems.
References:
"Linear Matrix Inequalities in System and Control Theory"
"Linear Controller Design: Limits of Performance" (both by Boyd).
References:
"Linear Matrix Inequalities in System and Control Theory"
"Linear Controller Design: Limits of Performance" (both by Boyd).
Sunday, April 27, 2014
Evan Archer: April 29th
Bayesian nonparametric methods for entropy estimation in spike data
Shannon’s entropy is a basic quantity in information theory, and a useful tool for the analysis of neural codes. However, estimating entropy from data is a difficult statistical problem. In this talk, I will discuss the problem of estimating entropy in the “under-sampled regime”, where the number of samples is small relative to the number of symbols. Dirichlet and Pitman-Yor processes provide tractable priors over countably-infinite discrete distributions, and have found applications in Bayesian non-parametric statistics and machine learning. In this talk, I will show that they also provide natural priors for Bayesian entropy estimation. These nonparametric priors permit us to address two major issues with previously-proposed Bayesian entropy estimators: their dependence on knowledge of the total number of symbols, and their inability to account for the heavy-tailed distributions which abound in biological and other natural data. What’s more, by “centering” a Dirichlet Process over a flexible parametric model, we are able to develop Bayesian estimators for the entropy of binary spike trains using priors designed to flexibly exploit the statistical structure of simultaneously-recorded spike responses. Finally, in applications to simulated and real neural data, I'll show that these estimators perform well in comparison to traditional methods.
Shannon’s entropy is a basic quantity in information theory, and a useful tool for the analysis of neural codes. However, estimating entropy from data is a difficult statistical problem. In this talk, I will discuss the problem of estimating entropy in the “under-sampled regime”, where the number of samples is small relative to the number of symbols. Dirichlet and Pitman-Yor processes provide tractable priors over countably-infinite discrete distributions, and have found applications in Bayesian non-parametric statistics and machine learning. In this talk, I will show that they also provide natural priors for Bayesian entropy estimation. These nonparametric priors permit us to address two major issues with previously-proposed Bayesian entropy estimators: their dependence on knowledge of the total number of symbols, and their inability to account for the heavy-tailed distributions which abound in biological and other natural data. What’s more, by “centering” a Dirichlet Process over a flexible parametric model, we are able to develop Bayesian estimators for the entropy of binary spike trains using priors designed to flexibly exploit the statistical structure of simultaneously-recorded spike responses. Finally, in applications to simulated and real neural data, I'll show that these estimators perform well in comparison to traditional methods.
Thursday, April 24, 2014
Maurizio Filippone: 30th April
Pseudo-Marginal Bayesian Inference for Gaussian Processes
Statistical models where parameters have a hierarchical structure are commonly employed to flexibly model complex phenomena and to gain some insight into the functioning of the system under study.
Carrying out exact parameter inference for such models, which is key to achieve a sound quantification of uncertainty in parameter estimates and predictions, usually poses a number of computational challenges. In this talk, I will focus on Markov chain Monte Carlo (MCMC) based inference for hierarchical models involving Gaussian Process (GP) priors and non-Gaussian likelihood functions.
After discussing why MCMC is the only way to infer parameters "exactly" in general GP models and pointing out the challenges in doing so, I will present a practical and efficient alternative to popular MCMC reparameterization techniques based on the so called Pseudo-Marginal MCMC approach.
In particular, the Pseudo-Marginal MCMC approach yields samples from the exact posterior distribution over GP covariance parameters, but only requires an unbiased estimate of the analytically intractable marginal likelihood. Finally, I will present ways to construct unbiased estimates of the marginal likelihood in GP models, and conclude the talk by presenting results on several benchmark data and on a multi-class multiple-kernel classification problem with neuroimaging data.
Useful links
http://www.dcs.gla.ac.uk/~maurizio/index.html
http://arxiv.org/abs/1310.0740
http://arxiv.org/abs/1311.7320
http://www.dcs.gla.ac.uk/~maurizio/Publications/aoas12.pdf
http://www.dcs.gla.ac.uk/~maurizio/Publications/ml13.pdf
Statistical models where parameters have a hierarchical structure are commonly employed to flexibly model complex phenomena and to gain some insight into the functioning of the system under study.
Carrying out exact parameter inference for such models, which is key to achieve a sound quantification of uncertainty in parameter estimates and predictions, usually poses a number of computational challenges. In this talk, I will focus on Markov chain Monte Carlo (MCMC) based inference for hierarchical models involving Gaussian Process (GP) priors and non-Gaussian likelihood functions.
After discussing why MCMC is the only way to infer parameters "exactly" in general GP models and pointing out the challenges in doing so, I will present a practical and efficient alternative to popular MCMC reparameterization techniques based on the so called Pseudo-Marginal MCMC approach.
In particular, the Pseudo-Marginal MCMC approach yields samples from the exact posterior distribution over GP covariance parameters, but only requires an unbiased estimate of the analytically intractable marginal likelihood. Finally, I will present ways to construct unbiased estimates of the marginal likelihood in GP models, and conclude the talk by presenting results on several benchmark data and on a multi-class multiple-kernel classification problem with neuroimaging data.
Useful links
http://www.dcs.gla.ac.uk/~maurizio/index.html
http://arxiv.org/abs/1310.0740
http://arxiv.org/abs/1311.7320
http://www.dcs.gla.ac.uk/~maurizio/Publications/aoas12.pdf
http://www.dcs.gla.ac.uk/~maurizio/Publications/ml13.pdf
Subscribe to:
Posts (Atom)