<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Arrykrishna Mootoovaloo</title>
    <description>Arrykrishna Mootoovaloo is a researcher in London, with a background in probabilistic machine learning and Bayesian inference from the University of Oxford.</description>
    <link>https://amootoovaloo.github.io/</link>
    <atom:link href="https://amootoovaloo.github.io/feed.xml" rel="self" type="application/rss+xml"/>
    <pubDate>Mon, 28 Sep 2026 15:46:31 +0000</pubDate>
    <lastBuildDate>Mon, 28 Sep 2026 15:46:31 +0000</lastBuildDate>
    <generator>Jekyll v4.4.1</generator>
    
      <item>
        <title>Hedging and Forward Curves in Energy Markets</title>
        <description>&lt;p align=&quot;justify&quot;&gt;An energy supplier agrees to deliver gas or electricity to its customers at a fixed price, often months or years in advance, but buys that energy in a wholesale market where prices move every day. If wholesale prices rise before the energy has been bought, the supplier&apos;s margin shrinks or disappears. Hedging, buying some of the energy ahead of time at today&apos;s prices, reduces that risk. In this post, I give a non-technical overview of two separate problems I have worked on in this area: deciding how much to hedge, and shaping a forward curve.&lt;/p&gt;

&lt;h2&gt;How much to hedge&lt;/h2&gt;

&lt;p align=&quot;justify&quot;&gt;The first problem is how much energy to buy in advance. The starting point is the &lt;i&gt;volume profile&lt;/i&gt;: how much energy the supplier expects its customers to consume in each future period. The question is then how much of that volume to hedge through the liquid contracts, those that can be traded readily at a fair price.&lt;/p&gt;

&lt;p align=&quot;justify&quot;&gt;Because future prices are uncertain, this decision is made by simulation. Using Monte Carlo methods, we generate many possible future paths for forward prices and for the spot price, the price for immediate delivery. Together with the volume profile, these simulations are all the optimisation needs. For a given choice of hedge volumes, we compute the supplier&apos;s total cost of supplying its customers along every path: the hedged volume is bought at forward prices, and the remainder is bought later at the spot price. This gives the full distribution of possible outcomes rather than a single forecast.&lt;/p&gt;

&lt;p align=&quot;justify&quot;&gt;The optimisation then searches for the hedge volumes that give the most favourable distribution, for example by reducing the chance of a very large cost. The appeal of simulation is that it captures the full range of possible price movements, rather than relying on a single forecast of where prices will go.&lt;/p&gt;

&lt;h2&gt;Forward curves&lt;/h2&gt;

&lt;p align=&quot;justify&quot;&gt;The second problem is a separate one. Energy for future delivery is traded through standard contracts, such as next month, next quarter, next season or next year, each with a single price covering its whole delivery period. A supplier, however, needs to know the price of energy for each individual day, or even each hour, because that is how its customers consume. A &lt;i&gt;forward curve&lt;/i&gt; fills this gap: it assigns a price to every day in the future while remaining consistent with the prices of the contracts that are actually traded.&lt;/p&gt;

&lt;p align=&quot;justify&quot;&gt;Building such a curve is known as forward shaping. A well-established approach is described by &lt;a href=&quot;https://ideas.repec.org/h/wsi/wschap/9789812812315_0007.html&quot;&gt;Benth, Šaltytė Benth and Koekebakker&lt;/a&gt;, originally for electricity markets. Among all the curves whose average over each contract&apos;s delivery period matches that contract&apos;s price, it chooses the smoothest one, so that prices do not jump artificially from one contract to the next. The curve can also follow a seasonal pattern, so that the known shape of demand through the year is reflected in the prices.&lt;/p&gt;

&lt;p align=&quot;justify&quot;&gt;The same idea carries over naturally to gas. The contracts differ, with gas markets trading months, quarters and the winter and summer seasons, and so does the seasonal pattern: gas demand is dominated by heating, so prices are typically highest in winter. With these adjustments, the method produces a daily gas curve that respects every traded price while varying smoothly from one day to the next.&lt;/p&gt;

&lt;h2&gt;Closing thoughts&lt;/h2&gt;

&lt;p align=&quot;justify&quot;&gt;The two problems are quite different. Hedging optimisation turns a vague question, &quot;how exposed are we?&quot;, into a concrete decision, by showing how different hedging choices would perform across a wide range of possible futures. Forward shaping provides a consistent view of prices for every future day from the handful of contracts that are actually traded. In both cases, as in many areas of applied statistics, the value lies less in predicting the future than in making the best use of the information available.&lt;/p&gt;
</description>
        <pubDate>Mon, 28 Sep 2026 08:00:00 +0000</pubDate>
        <link>https://amootoovaloo.github.io/blog/2026/09/Hedging-Energy-Forward-Curves</link>
        <guid isPermaLink="true">https://amootoovaloo.github.io/blog/2026/09/Hedging-Energy-Forward-Curves</guid>
        
        
        <category>Machine Learning and Statistics</category>
        
      </item>
    
      <item>
        <title>Learning Momentum Across Time and Assets</title>
        <description>&lt;p align=&quot;justify&quot;&gt;Over the past few months, I have been studying how machine learning can be used to build systematic trading strategies, and in particular how to make them robust once trading costs are taken into account. A useful reference point is the paper &lt;a href=&quot;https://arxiv.org/abs/2302.10175&quot;&gt;Spatio-Temporal Momentum&lt;/a&gt; by Tan, Roberts and Zohren. In this post, I give a non-technical overview of its idea, followed by two topics that matter as much as the model itself: controlling turnover and backtesting honestly.&lt;/p&gt;

&lt;h2&gt;Momentum&lt;/h2&gt;

&lt;p align=&quot;justify&quot;&gt;Momentum is one of the most widely documented patterns in financial markets: assets that have performed well recently tend to continue performing well for a while, and those that have performed poorly tend to continue to lag. It comes in two main forms.&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;p align=&quot;justify&quot;&gt;&lt;b&gt;Time-series momentum&lt;/b&gt; looks at each asset on its own. If an asset&apos;s price has been rising, we buy it; if it has been falling, we sell it.&lt;/p&gt;&lt;/li&gt;
  &lt;li&gt;&lt;p align=&quot;justify&quot;&gt;&lt;b&gt;Cross-sectional momentum&lt;/b&gt; compares assets with one another. We rank them by recent performance, buy the strongest and sell the weakest, regardless of whether the market as a whole is rising or falling.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p align=&quot;justify&quot;&gt;Traditionally, these two strategies are designed and studied separately, even though they draw on the same information.&lt;/p&gt;

&lt;h2&gt;The idea&lt;/h2&gt;

&lt;p align=&quot;justify&quot;&gt;The paper&apos;s central idea is to learn both at once. Rather than deciding each asset&apos;s position from its own history alone, the model looks at the recent behaviour of all the assets together and produces positions for all of them in one step. It can therefore learn, for example, that a move in one asset carries information about another.&lt;/p&gt;

&lt;p align=&quot;justify&quot;&gt;Perhaps the most striking finding is how simple the model can be. The authors show that a neural network with a single layer, essentially a learned weighting of the input features, is enough to generate useful signals for every asset simultaneously. The approach was tested on US equities and on equity index futures, and it remained competitive with standard benchmarks once realistic transaction costs were included.&lt;/p&gt;

&lt;h2&gt;Turnover control&lt;/h2&gt;

&lt;p align=&quot;justify&quot;&gt;A strategy that looks profitable on paper can lose money in practice if it trades too often, because every trade incurs costs. &lt;i&gt;Turnover&lt;/i&gt; measures how much of the portfolio changes from one period to the next. A good way to keep it in check is to build it into training: one penalty encourages smaller positions unless the evidence is strong, and another discourages large changes in positions from one day to the next. The paper finds that combining the two gives the best results once costs are included.&lt;/p&gt;

&lt;h2&gt;Backtesting honestly&lt;/h2&gt;

&lt;p align=&quot;justify&quot;&gt;A backtest simulates how a strategy would have performed on historical data, and it is easy to be fooled by one. The essentials are to use only information that was available at the time, to evaluate on a later period than the model was trained on, to include assets that were delisted or failed, to report results after realistic costs, and to remember that if enough ideas are tried, some will look good purely by chance.&lt;/p&gt;

&lt;h2&gt;Closing thoughts&lt;/h2&gt;

&lt;p align=&quot;justify&quot;&gt;What I find most appealing about this line of work is its restraint. A simple model that shares information across assets, trained with an explicit awareness of trading costs and evaluated with a careful backtest, is often more valuable than a complex model evaluated optimistically. In quantitative research, how a strategy is tested is as important as how it is built.&lt;/p&gt;
</description>
        <pubDate>Sat, 15 Nov 2025 08:00:00 +0000</pubDate>
        <link>https://amootoovaloo.github.io/blog/2025/11/Learning-Momentum-Across-Assets</link>
        <guid isPermaLink="true">https://amootoovaloo.github.io/blog/2025/11/Learning-Momentum-Across-Assets</guid>
        
        
        <category>Machine Learning and Statistics</category>
        
      </item>
    
      <item>
        <title>An Introduction to On-Device AI</title>
        <description>&lt;p align=&quot;justify&quot;&gt;This is a short summary of the course &lt;a href=&quot;https://learn.deeplearning.ai/courses/introduction-to-on-device-ai/lesson/1/introduction&quot;&gt;&lt;strong&gt;Introduction to on-device AI&lt;/strong&gt;&lt;/a&gt; provided by DeepLearning.ai.&lt;/p&gt;

&lt;section&gt;
    &lt;h1&gt;Introduction&lt;/h1&gt;
    &lt;p align=&quot;justify&quot;&gt;Modern smartphones, with their significant computational power of 10 to 30 teraflops, can run multiple AI models simultaneously for applications like real-time semantic segmentation and scene understanding. This course explores creating AI applications that run on-device, applicable to smartphones and billions of other devices such as cameras, robots, drones, and VR headsets. Despite differences in hardware and operating systems, the principles for deploying on-device AI are similar. First, models trained in the cloud must be converted into a compatible format, which involves freezing the model into a neural network graph and converting it into an executable format for the target device.&lt;/p&gt;
    &lt;p align=&quot;justify&quot;&gt;Devices like smartphones and edge devices usually contain a mix of CPUs, GPUs, and NPUs, enabling performance optimizations tailored to specific hardware. Understanding the target devices can significantly enhance performance, making applications run up to ten times faster. The course also covers tools for ensuring consistent model performance across devices, which is crucial given the diversity of more than 300 smartphone types. Validating numerical correctness prevents discrepancies in model behaviour caused by hardware differences. Quantization, which reduces the numerical precision of the model&apos;s calculations, is another important aspect: it can make apps run several times faster while reducing model size and memory footprint.&lt;/p&gt;
&lt;/section&gt;

&lt;section&gt;
    &lt;h1&gt;Why On-Device&lt;/h1&gt;
    &lt;p align=&quot;justify&quot;&gt;On-device AI offers several benefits, including reduced latency, improved efficiency, lower costs, and enhanced privacy. Applications of on-device AI are prevalent in real-world scenarios such as real-time speech detection, semantic segmentation, object detection, and physical activity detection. Each time you take a picture with your smartphone, over 20 AI models work in tandem to optimize the image within milliseconds. The industrial IoT sector alone sees an estimated economic impact of around $3 trillion due to these technologies.&lt;/p&gt;
    &lt;p align=&quot;justify&quot;&gt;On-device AI is integrated into many everyday technologies. For instance, typing on a laptop keyboard or interacting with a smart speaker relies on AI running locally. Robots for delivery and assembly, as well as drones used in industrial and agricultural settings, utilize on-device AI. This technology is also pivotal in advanced driver assistance systems in cars and in editing photos on smartphones and laptops. Various applications across audio, image, video, and sensor data—like text-to-speech, speech recognition, photo classification, and physical activity detection—demonstrate the versatility of on-device AI. Running models on-device is advantageous for several reasons:&lt;/p&gt;
    &lt;ul&gt;
        &lt;li&gt;It is cost-effective as it leverages local computational resources, thereby eliminating the need for cloud services.&lt;/li&gt;
        &lt;li&gt;It is more efficient, as data are processed locally, avoiding the delays associated with cloud communication.&lt;/li&gt;
        &lt;li&gt;Privacy is enhanced since data remains on the device.&lt;/li&gt;
        &lt;li&gt;It enables personalization, as models can be tailored locally without external data.&lt;/li&gt;
    &lt;/ul&gt;
&lt;/section&gt;

&lt;section&gt;
    &lt;h1&gt;Deploying Segmentation Models On-Device&lt;/h1&gt;
    &lt;p align=&quot;justify&quot;&gt;On-device AI has a wide range of applications, including real-time object detection to identify people, faces, and QR codes; speech recognition to convert spoken input into text; and pose estimation to predict human poses from real-time images or video streams. Other applications include image generation from text descriptions, super-resolution to enhance low-resolution images, and image segmentation, which divides an image into meaningful segments for easier analysis and object identification.&lt;/p&gt;
    &lt;p align=&quot;justify&quot;&gt;Image segmentation comes in various forms, with the most popular being semantic segmentation and instance segmentation.&lt;/p&gt;
    &lt;ul&gt;
        &lt;li&gt;Semantic segmentation labels every pixel according to a specific class.&lt;/li&gt;
        &lt;li&gt;Instance segmentation labels each pixel by class and individual instance within the same category.&lt;/li&gt;
    &lt;/ul&gt;
    &lt;p align=&quot;justify&quot;&gt;The latter is widely used in advanced driver assistance systems to differentiate roads from pedestrians, in image editing software to apply filters or blur backgrounds, and in video conferencing to blur backgrounds during calls. It is also employed in drones for mapping landscapes in agricultural and industrial settings.&lt;/p&gt;
    &lt;p align=&quot;justify&quot;&gt;Several models are used for semantic segmentation, including ResNet, which employs residual connections for training deep networks, and HRNet, which captures fine-grained details through high-resolution representations. FANet, or Feature Agglomeration Network, aggregates features from different scales for detailed predictions, while DDRNet, or Dual Dynamic Resolution Network, balances efficiency and accuracy with a dual-path architecture.&lt;/p&gt;
    &lt;p align=&quot;justify&quot;&gt;The course focuses on FFNET, or Fuss Free Network, which features a simple encoder-decoder architecture with a ResNet-like backbone and a small multi-scale head. FFNET offers comparable accuracy to more complex networks like HRNet and FANet, but with greater computational efficiency, making it ideal for on-device deployment. Its configurable design supports a range of encoder and decoder sizes to suit different deployment environments and accuracy requirements.&lt;/p&gt;
&lt;/section&gt;

&lt;section&gt;
    &lt;h1&gt;Preparing for On-Device Deployment&lt;/h1&gt;
    &lt;p align=&quot;justify&quot;&gt;The first step is neural network graph capture, which involves capturing the computation to be executed on the device. The second step is on-device compilation. The third step focuses on accelerating models using the device&apos;s hardware. The fourth step emphasizes validating the numerical correctness of the model on the device.&lt;/p&gt;
    &lt;p align=&quot;justify&quot;&gt;There are three popular runtimes for on-device deployment: TensorFlow Lite, recommended for Android applications; ONNX Runtime, suitable for Windows-based applications; and the Qualcomm AI Engine, ideal for fully embedded applications on Qualcomm hardware.&lt;/p&gt;
    &lt;p align=&quot;justify&quot;&gt;We now look at the TensorFlow Lite runtime in more detail. Designed specifically for mobile platforms and embedded devices, TensorFlow Lite delivers highly efficient performance and fast response times by minimizing computational overhead. Its flexibility allows deployment on smartphones and a wide range of IoT devices, making it highly portable. TensorFlow Lite is also energy-efficient. It supports hardware acceleration and is fully compatible with neural processing units (NPUs) through a mechanism called delegation.&lt;/p&gt;
&lt;/section&gt;

&lt;section&gt;
    &lt;h1&gt;Quantizing Models&lt;/h1&gt;
    &lt;p align=&quot;justify&quot;&gt;Quantization reduces the precision of your model to speed up computation and decrease model size. This process can make models up to four times smaller and faster. The three main benefits of quantization are: reduced model size for better storage on devices with limited capacity, faster processing due to fewer computations, and lower power consumption, which is crucial for battery-operated devices.&lt;/p&gt;
    &lt;p align=&quot;justify&quot;&gt;To understand quantization, consider a floating point tensor, which takes 32 bits per value. Quantization converts this to an integer representation with 8 bits per value, using a scale and zero point for accurate translation. The goal is to minimize the error introduced when converting from floating point to integer.&lt;/p&gt;
    &lt;p align=&quot;justify&quot;&gt;There are two main types of quantization:&lt;/p&gt;
    &lt;ul&gt;
        &lt;li&gt;Weight quantization, which reduces the precision of model weights to optimize storage.&lt;/li&gt;
        &lt;li&gt;Activation quantization, which applies lower precision to activation values to accelerate inference using lower precision numerics.&lt;/li&gt;
    &lt;/ul&gt;
    &lt;p align=&quot;justify&quot;&gt;Different levels of quantization exist, such as:&lt;/p&gt;
    &lt;ul&gt;
        &lt;li&gt;W8int8, converting both weights and activations to 8-bit.&lt;/li&gt;
        &lt;li&gt;W8int16, converting weights to 8-bit and keeping activations at 16-bit.&lt;/li&gt;
        &lt;li&gt;Extreme quantization like W4 (weights at 4-bit and activations at 16-bit) is popular for large language models and generative AI applications.&lt;/li&gt;
    &lt;/ul&gt;
    &lt;p align=&quot;justify&quot;&gt;Quantization methods include:&lt;/p&gt;
    &lt;ul&gt;
        &lt;li&gt;Post-training quantization&lt;/li&gt;
        &lt;li&gt;Quantization-aware training&lt;/li&gt;
    &lt;/ul&gt;
    &lt;p align=&quot;justify&quot;&gt;Post-training quantization applies the process after training, using calibration with sample data to minimize quantization error. This typically requires a few hundred samples.&lt;/p&gt;
    &lt;p align=&quot;justify&quot;&gt;Quantization-aware training incorporates quantization into the training process itself, learning the model weights and their integer representation simultaneously. This can yield more accurate models when post-training quantization is insufficient.&lt;/p&gt;
&lt;/section&gt;

&lt;section&gt;
    &lt;h1&gt;Device Integration&lt;/h1&gt;
    &lt;p align=&quot;justify&quot;&gt;The following outlines the steps for integrating an AI model into a smartphone application for real-time segmentation at 30 frames per second:&lt;/p&gt;
    &lt;ol align=&quot;justify&quot;&gt;
        &lt;li&gt;&lt;strong&gt;Understanding Data Transformation&lt;/strong&gt;: The data flow begins with the camera stream, which provides frames in either RGB or YUV format. The GPU is used for pre-processing to convert YUV to RGB and downsample to the model&apos;s required resolution (e.g., from 720p to 224x224).&lt;/li&gt;
        &lt;li&gt;&lt;strong&gt;Device Integration Process&lt;/strong&gt;: Integration involves several stages:
            &lt;ul&gt;
                &lt;li&gt;Camera Stream Extraction: Extract frames at 30 frames per second in either RGB or YUV format.&lt;/li&gt;
                &lt;li&gt;Pre-processing: Utilize the GPU and OpenCV for efficient conversion and downsampling of the frames to match the model&apos;s resolution.&lt;/li&gt;
                &lt;li&gt;Model Inference: Run the AI model on the device&apos;s Neural Processing Unit (NPU) using runtime APIs (C++ or Java-based) for optimal performance. The model outputs a mask of the same resolution as it was trained on (e.g., 224x224).&lt;/li&gt;
                &lt;li&gt;Post-processing: Use the GPU and OpenCV again for tasks like upsampling, smoothing, thresholding, and blurring to overlay the output mask onto the camera stream.&lt;/li&gt;
                &lt;li&gt;Packaging Runtime: Ensure all runtime dependencies, including the AI models, are packaged within the application to leverage hardware acceleration effectively.&lt;/li&gt;
            &lt;/ul&gt;
        &lt;/li&gt;
        &lt;li&gt;&lt;strong&gt;Implementation&lt;/strong&gt;: There are five key parts:
            &lt;ul&gt;
                &lt;li&gt;Camera Stream Extraction: Capture frames from the camera.&lt;/li&gt;
                &lt;li&gt;Pre-processing: Convert and downsample frames using the GPU.&lt;/li&gt;
                &lt;li&gt;Model Inference: Perform inference on the NPU.&lt;/li&gt;
                &lt;li&gt;Post-processing: Overlay the model&apos;s output on the camera stream using GPU-based techniques.&lt;/li&gt;
                &lt;li&gt;Packaging: Include all necessary runtime dependencies in the application to ensure it runs efficiently on the device. This involves managing Java, native sources, and dependencies within an Android project.&lt;/li&gt;
            &lt;/ul&gt;
        &lt;/li&gt;
    &lt;/ol&gt;
&lt;/section&gt;
</description>
        <pubDate>Sat, 08 Jun 2024 08:00:00 +0000</pubDate>
        <link>https://amootoovaloo.github.io/blog/2024/06/Introduction-to-On-Device</link>
        <guid isPermaLink="true">https://amootoovaloo.github.io/blog/2024/06/Introduction-to-On-Device</guid>
        
        
        <category>Machine Learning and Statistics</category>
        
      </item>
    
      <item>
        <title>emuflow: Combining Experiments with Normalising Flows</title>
        <description>&lt;p align=&quot;justify&quot;&gt;In this post, I briefly summarise &lt;a href=&quot;https://arxiv.org/abs/2409.01407&quot;&gt;emuflow&lt;/a&gt;, written at Oxford with Carlos García-García, David Alonso and Jaime Ruiz-Zapatero. The question is simple: if several experiments have already published MCMC chains, can we combine them without re-running their expensive likelihoods?&lt;/p&gt;

&lt;h2&gt;The problem&lt;/h2&gt;

&lt;p align=&quot;justify&quot;&gt;Each experiment constrains a small set of shared cosmological parameters $\boldsymbol{\theta}$, but also carries its own nuisance parameters $\boldsymbol{\beta}_i$ (galaxy biases, calibration, redshift shifts and so on). With $b$ cosmological parameters and $c_i$ nuisance parameters per experiment, a joint analysis samples $b+\sum_i c_i$ dimensions, and every step calls every experiment&apos;s forward model. The dimension and the cost both grow with each dataset added.&lt;/p&gt;

&lt;p align=&quot;justify&quot;&gt;Yet most of the time, we only care about $p(\boldsymbol{\theta}\,|\,\boldsymbol{x}_i)$, the posterior with the nuisance parameters marginalised out. Existing chains already contain samples from it: we just drop the nuisance columns. What is missing is a density we can evaluate.&lt;/p&gt;

&lt;h2&gt;The idea&lt;/h2&gt;

&lt;p align=&quot;justify&quot;&gt;A normalising flow starts from a simple distribution, such as a Gaussian, and passes it through a sequence of invertible neural network transformations that stretch and bend it into the shape of the target posterior. Because every step is invertible, the flow can both generate new samples and evaluate the density at any point, which is exactly what we need. Training is maximum likelihood on the chain samples. We use an affine autoregressive flow, which keeps the density calculation cheap. About 20,000 samples train a flow over five or six parameters in roughly two minutes on a desktop.&lt;/p&gt;

&lt;p align=&quot;justify&quot;&gt;Once each experiment has a flow, there are two ways to use it. A legacy experiment&apos;s flow can act as an informative prior for the likelihood of a new one, so the old nuisance parameters and forward model disappear from the analysis. Alternatively, the flows alone can be multiplied together (a product of experts), correcting for the shared prior:&lt;/p&gt;

\[p(\boldsymbol{\theta}\,|\,\boldsymbol{x}_1,\ldots,\boldsymbol{x}_N) \propto p(\boldsymbol{\theta}) \prod_{i=1}^{N} \frac{p_{\textrm{nf}}(\boldsymbol{\theta}\,|\,\boldsymbol{x}_i)}{p(\boldsymbol{\theta})}.\]

&lt;h2&gt;A hard test&lt;/h2&gt;

&lt;p align=&quot;justify&quot;&gt;We validated the method on a deliberately difficult pair: Planck 2018 and a large combination of large-scale structure data (galaxy clustering, weak lensing and CMB lensing), which disagree on $S_8$ at about $3.5\sigma$. Their joint posterior sits in the tails of each individual one, so a flow that only captures the bulk would fail.&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;b&gt;Planck flow as a prior:&lt;/b&gt; posterior means are recovered at the sub-percent level and widths within 2–3%. The exact joint run took about 24 days on two HPC nodes; this took about 6.&lt;/li&gt;
  &lt;li&gt;&lt;b&gt;Two flows only:&lt;/b&gt; the joint posterior takes under 15 minutes on a desktop. The price is accuracy: means shift by up to about $0.3\sigma$, which is expected given how much the result relies on the tails.&lt;/li&gt;
&lt;/ul&gt;

&lt;figure class=&quot;figure aligncenter&quot; style=&quot;max-width: 700px&quot;&gt;
  &lt;img src=&quot;/images/blog/emuflow/posterior.webp&quot; alt=&quot;Joint constraints from the large-scale structure data and Planck. Using the Planck flow as a prior (green) recovers the full joint analysis (dark blue).&quot; loading=&quot;lazy&quot; /&gt;
  &lt;figcaption&gt;Joint constraints from the large-scale structure data and Planck. Using the Planck flow as a prior (green) recovers the full joint analysis (dark blue).&lt;/figcaption&gt;
&lt;/figure&gt;

&lt;p align=&quot;justify&quot;&gt;Further tests, combining KiDS-1000 with DES Y3 and Planck with DES Y1, involve milder tension but less Gaussian posteriors, and lead to the same conclusions.&lt;/p&gt;

&lt;h2&gt;A zoo of pre-trained flows&lt;/h2&gt;

&lt;p align=&quot;justify&quot;&gt;We also trained flows on public chains from Planck 2018, DES Y3, KiDS-1000, ACT DR4 and SDSS, and released them with the &lt;a href=&quot;https://github.com/Harry45/emuflow/&quot;&gt;code&lt;/a&gt;. Adding a new experiment or extending the cosmological model only needs its MCMC samples and a few minutes of training.&lt;/p&gt;

&lt;p align=&quot;justify&quot;&gt;One caveat: the combination rule assumes independent datasets and uninformative priors on $\boldsymbol{\theta}$ in each original analysis. A chain that already includes another experiment as a prior would count that information twice.&lt;/p&gt;
</description>
        <pubDate>Mon, 15 Jan 2024 08:00:00 +0000</pubDate>
        <link>https://amootoovaloo.github.io/blog/2024/01/emuflow-Normalising-Flows</link>
        <guid isPermaLink="true">https://amootoovaloo.github.io/blog/2024/01/emuflow-Normalising-Flows</guid>
        
        
        <category>Machine Learning and Statistics</category>
        
      </item>
    
      <item>
        <title>Editing Images by Editing Sentences</title>
        <description>&lt;p align=&quot;justify&quot;&gt;Over the past few months, I have been exploring sequential image editing: making a series of changes to an image, one after another, by describing each change in words. In this post, I give a non-technical overview of the idea and of two papers that shaped it.&lt;/p&gt;

&lt;h2&gt;The goal&lt;/h2&gt;

&lt;p align=&quot;justify&quot;&gt;Text-to-image models can turn a sentence such as &lt;i&gt;&quot;a basket of apples&quot;&lt;/i&gt; into a convincing picture. Editing that picture is harder. If we simply ask the model for &lt;i&gt;&quot;a basket of eggs&quot;&lt;/i&gt;, we get a new image altogether: a different basket, a different table, different lighting. What we usually want is the same picture with only the apples replaced by eggs, and then, perhaps, a second and third change on top of that.&lt;/p&gt;

&lt;h2&gt;How diffusion models create images&lt;/h2&gt;

&lt;p align=&quot;justify&quot;&gt;Diffusion models learn to generate images by learning to remove noise. During training, noise is gradually added to real images until nothing recognisable remains, and the model learns to reverse each step. To create a new image, the model starts from pure noise and removes it step by step until a clean picture emerges.&lt;/p&gt;

&lt;p align=&quot;justify&quot;&gt;When a sentence is provided, it is first converted into a numerical representation, a list of numbers that captures its meaning. This representation guides every denoising step, steering the emerging image towards the description. Much of an image&apos;s overall layout is decided in the early, noisy steps, while the later steps fill in finer details.&lt;/p&gt;

&lt;h2&gt;Two reference papers&lt;/h2&gt;

&lt;p align=&quot;justify&quot;&gt;&lt;a href=&quot;https://arxiv.org/abs/2108.01073&quot;&gt;SDEdit&lt;/a&gt; (Meng et al.) showed that an existing image can be edited without retraining the model. Instead of starting from pure noise, we add a moderate amount of noise to the image we already have, then let the model remove it again. Because the noise does not erase everything, the broad structure of the original survives, while the model is free to redraw the details. The amount of noise sets the balance: too little and nothing changes, too much and the original is lost.&lt;/p&gt;

&lt;p align=&quot;justify&quot;&gt;&lt;a href=&quot;https://arxiv.org/abs/2208.01626&quot;&gt;Prompt-to-Prompt&lt;/a&gt; (Hertz et al.) looked at how each word in a sentence influences each part of the image. Inside the model, a mechanism called cross-attention links every word to the regions of the image it affects, so the word &lt;i&gt;&quot;apples&quot;&lt;/i&gt; is tied mainly to the pixels where the apples appear. The authors showed that if these links are kept from the original image while one word in the sentence is changed, the layout and composition are preserved and only the relevant region changes.&lt;/p&gt;

&lt;h2&gt;The idea&lt;/h2&gt;

&lt;p align=&quot;justify&quot;&gt;The approach I have been exploring brings these two ideas together. An image is generated, or taken as a starting point, together with the sentence that describes it, such as &lt;i&gt;&quot;a basket of apples&quot;&lt;/i&gt;. To edit it, we write the new description, &lt;i&gt;&quot;a basket of eggs&quot;&lt;/i&gt;, and replace the representation of the original sentence with that of the new one inside the diffusion model. Starting from the original image rather than from scratch, the model then denoises under the guidance of the new sentence. Because only one word has changed, the two representations are very similar, so the model keeps the basket, the table and the lighting, and redraws only what the new word requires.&lt;/p&gt;

&lt;p align=&quot;justify&quot;&gt;The same step can be repeated, which is what makes the editing sequential. The edited image and its new sentence become the starting point for the next change: &lt;i&gt;&quot;a basket of eggs&quot;&lt;/i&gt; can become &lt;i&gt;&quot;a wicker basket of eggs&quot;&lt;/i&gt;, and then &lt;i&gt;&quot;a wicker basket of eggs on a wooden table&quot;&lt;/i&gt;. Each edit is expressed in plain language, and each builds on the result of the previous one rather than starting again.&lt;/p&gt;

&lt;h2&gt;Challenges&lt;/h2&gt;

&lt;p align=&quot;justify&quot;&gt;Sequential editing brings its own difficulties. Small imperfections can accumulate over several edits, so an image may gradually drift away from the original. Some changes are also harder than others: replacing one object with another of similar size and shape, such as apples with eggs, is far easier than a change that alters the whole composition. Finally, editing a real photograph, rather than an image the model generated itself, first requires finding the noise and sentence that would reproduce it, which is an active research problem in its own right.&lt;/p&gt;

&lt;p align=&quot;justify&quot;&gt;What makes this line of work appealing is that the interface is language. Rather than masking regions or adjusting sliders, a user describes the change they want, and the model works out where and how to apply it.&lt;/p&gt;
</description>
        <pubDate>Sat, 15 Jul 2023 08:00:00 +0000</pubDate>
        <link>https://amootoovaloo.github.io/blog/2023/07/Editing-Images-by-Editing-Sentences</link>
        <guid isPermaLink="true">https://amootoovaloo.github.io/blog/2023/07/Editing-Images-by-Editing-Sentences</guid>
        
        
        <category>Machine Learning and Statistics</category>
        
      </item>
    
      <item>
        <title>An Emulator for the 3D Matter Power Spectrum</title>
        <description>&lt;p align=&quot;justify&quot;&gt;The 3D matter power spectrum, which describes how matter clumps on different scales and at different times, is a key quantity that underpins most cosmological data analyses, including galaxy clustering, weak lensing and 21 cm cosmology. Crucially, other (derived) power spectra can be calculated quickly once it has been precomputed. In practice, the matter power spectrum is the most expensive component: it is calculated either with Boltzmann solvers such as CLASS or CAMB, or with simulations, which can be computationally expensive depending on the resolution required.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/images/blog/power-spectrum-emulator/power-spectrum.webp&quot; alt=&quot;The 3D matter power spectrum as a function of wavenumber and redshift&quot; align=&quot;left&quot; width=&quot;400&quot; style=&quot;margin-right: 10px; margin-bottom: 10px&quot; /&gt;&lt;/p&gt;

&lt;p align=&quot;justify&quot;&gt;This work was published in &lt;a href=&quot;https://doi.org/10.1016/j.ascom.2021.100508&quot;&gt;Astronomy and Computing&lt;/a&gt;. Our contributions are threefold. First, we show that emulation does not always require a zero-mean Gaussian Process; additional basis functions can be included before defining the kernel matrix. This is useful when an approximate model of the function is already available. Moreover, if we know how a particular function behaves, we can adopt a stringent prior on the regression coefficients of the parametric model, encoding our degree of belief in that model. Second, because the Radial Basis Function (RBF) kernel we use is infinitely differentiable, we can estimate the first and second derivatives of the 3D matter power spectrum. The derived expressions for the derivatives involve only element-wise matrix multiplication, with no matrix inverse to compute, making the gradient calculations very fast. Finally, we show that the emulator can output several key power spectra: the linear matter power spectrum at a reference redshift, and the non-linear 3D matter power spectrum with or without an analytic baryon feedback model. Using the emulated 3D power spectrum together with the tomographic redshift distributions, we also show that the weak lensing and intrinsic alignment (II and GI) power spectra can be generated very quickly using existing numerical techniques. To make the problem tractable, the 3D matter power spectrum is split into three simpler pieces: the linear power spectrum at a reference redshift, a growth factor describing how it evolves with redshift, and a correction for non-linear effects on small scales. Each piece is modelled by its own semi-parametric Gaussian Process, with a second-order polynomial for the parametric part.&lt;/p&gt;

&lt;figure class=&quot;figure aligncenter&quot; style=&quot;max-width: 800px&quot;&gt;
  &lt;img src=&quot;/images/blog/power-spectrum-emulator/gradients.webp&quot; alt=&quot;The gradients of the power spectrum with respect to each cosmological parameter at a fixed redshift.&quot; loading=&quot;lazy&quot; /&gt;
  &lt;figcaption&gt;The gradients of the power spectrum with respect to each cosmological parameter at a fixed redshift.&lt;/figcaption&gt;
&lt;/figure&gt;

&lt;p align=&quot;justify&quot;&gt;The figure above shows the gradients at a fixed set of cosmological parameters (a test point) and a fixed redshift, $z=0$. The red curves show the gradients calculated by CLASS using the central difference method, and the blue curves show those output by the emulator. In general, the emulator returns the gradient for every combination of scale, redshift and cosmological parameter at once: by default, 1000 scales, 100 redshifts and 5 parameters.&lt;/p&gt;

&lt;figure class=&quot;figure aligncenter&quot; style=&quot;max-width: 600px&quot;&gt;
  &lt;img src=&quot;/images/blog/power-spectrum-emulator/posterior.webp&quot; alt=&quot;The full posterior distribution of all parameters using the emulator on a toy dataset.&quot; loading=&quot;lazy&quot; /&gt;
  &lt;figcaption&gt;The full posterior distribution of all parameters using the emulator on a toy dataset.&lt;/figcaption&gt;
&lt;/figure&gt;

&lt;p align=&quot;justify&quot;&gt;We also tested the emulator on simulated weak-lensing bandpowers. The mock survey has five redshift bins and ten bandpowers per pair of bins, giving 150 data points, with simple independent Gaussian errors of 50%. For simplicity, intrinsic alignments are switched off, although they can easily be included and marginalised over. The cosmological parameters used to generate the data are shown by the black dots in the figure above. We use a Gaussian likelihood and uniform priors on all cosmological parameters, matching the input range of the emulator. The figure above shows the results of sampling the cosmological parameters on this toy data set: the red contours correspond to the emulator, and the pale blue contours to the posterior distributions obtained with CLASS. We ran three separate MCMC chains of 150 000 samples each, two with the emulator and one with CLASS, and computed the Gelman-Rubin convergence statistic for each of the three resulting pairs of runs. The worst value is 1.002, consistent with all three chains being drawn from the same distribution and corroborating the agreement shown in the figure. The emulator developed in this work therefore robustly recovers the posterior distributions of all the cosmological parameters, in agreement with the accurate solver, CLASS.&lt;/p&gt;

</description>
        <pubDate>Tue, 16 Nov 2021 07:11:00 +0000</pubDate>
        <link>https://amootoovaloo.github.io/blog/2021/11/Imperial-Publication-2</link>
        <guid isPermaLink="true">https://amootoovaloo.github.io/blog/2021/11/Imperial-Publication-2</guid>
        
        
        <category>Machine Learning and Statistics</category>
        
      </item>
    
      <item>
        <title>Parameter Inference with MOPED and Gaussian Processes</title>
        <description>&lt;p align=&quot;justify&quot;&gt;In this post, I briefly summarise my first PhD paper, published in &lt;a href=&quot;https://academic.oup.com/mnras/article/497/2/2213/5873022&quot;&gt;MNRAS&lt;/a&gt;. A key step in this work is the compression and emulation of the data using MOPED, an algorithm developed by &lt;a href=&quot;https://academic.oup.com/mnras/article/317/4/965/1039456&quot;&gt;Heavens et al. 2000&lt;/a&gt;.&lt;/p&gt;

&lt;p align=&quot;justify&quot;&gt;A weak lensing data vector can contain hundreds or thousands of numbers, yet the model it constrains has only a handful of parameters. MOPED compresses a data vector of size $N$ to just $p$ numbers, one per parameter. Each compressed number is a weighted sum of the data, with weights chosen so that it captures as much information as possible about its parameter. The weights are built from the derivatives of the model with respect to each parameter, scaled by the inverse of the noise covariance, and are made mutually orthogonal and normalised. For Gaussian data whose noise covariance does not depend on the parameters, this compression loses no information at the fiducial parameters used to construct it.&lt;/p&gt;

&lt;p align=&quot;justify&quot;&gt;The compression also makes the likelihood remarkably simple. Because the weighting vectors are orthogonal and normalised, the compressed numbers are uncorrelated with unit variance, so the log-likelihood reduces to a sum of $p$ squared differences:&lt;/p&gt;

\[\log\mathcal{L} = -\frac{1}{2}\sum_{\alpha=1}^{p}\left(y_{\alpha} - \langle y_{\alpha}\rangle\right)^{2} + \textrm{constant}\]

&lt;p align=&quot;justify&quot;&gt;where $y_{\alpha}$ are the compressed data and $\langle y_{\alpha}\rangle$ their theoretical predictions. An MCMC algorithm can then sample the posterior distribution of the model parameters directly from the compressed data.&lt;/p&gt;

&lt;p align=&quot;justify&quot;&gt;However, computing the theoretical predictions at each step of an MCMC can still be costly when the forward model itself is expensive. We therefore first generate a training set of $N$ Latin Hypercube samples (LHS), compute the MOPED coefficients at these points, and then model them with $p$ separate Gaussian Processes. These serve as surrogates for sampling the posterior distribution of the model parameters, with the result shown in the figure below: the posterior obtained with the full, accurate solver CLASS is shown in tan, and the posterior obtained with the emulator in blue. The contours correspond to the 68% and 95% credible intervals.&lt;/p&gt;

&lt;figure class=&quot;figure aligncenter&quot; style=&quot;max-width: 800px&quot;&gt;
  &lt;img src=&quot;/images/blog/moped-gp/posterior.webp&quot; alt=&quot;The full posterior distribution of all parameters using the MOPED compression scheme.&quot; loading=&quot;lazy&quot; /&gt;
  &lt;figcaption&gt;The full posterior distribution of all parameters using the MOPED compression scheme.&lt;/figcaption&gt;
&lt;/figure&gt;

</description>
        <pubDate>Tue, 16 Nov 2021 07:11:00 +0000</pubDate>
        <link>https://amootoovaloo.github.io/blog/2021/11/Imperial-Publication-1</link>
        <guid isPermaLink="true">https://amootoovaloo.github.io/blog/2021/11/Imperial-Publication-1</guid>
        
        
        <category>Machine Learning and Statistics</category>
        
      </item>
    
      <item>
        <title>Machine Learning Resources</title>
        <description>&lt;p align=&quot;justify&quot;&gt;Below is a curated list of courses, lecture series and schools that I regularly use as references for learning machine learning techniques. Deep learning itself spans several branches, such as natural language processing and meta-learning; resources for these are listed separately.&lt;/p&gt;

&lt;p&gt;&lt;b&gt;General Machine Learning&lt;/b&gt;&lt;/p&gt;

&lt;ol&gt;

&lt;li&gt; The AI Epiphany (&lt;a href=&quot;https://www.youtube.com/channel/UCj8shE7aIn4Yawwbo2FceCQ&quot;&gt;Videos&lt;/a&gt;)&lt;/li&gt;

&lt;li&gt; ML Tech Talks (&lt;a href=&quot;https://www.youtube.com/playlist?list=PLQY2H8rRoyvwmjfn7hM-Yg_6RIyoMnKQx&quot;&gt;Videos&lt;/a&gt;)&lt;/li&gt;

&lt;li&gt;CPSC 540: Machine Learning 2013 (&lt;a href=&quot;https://www.cs.ubc.ca/~nando/540-2013/lectures.html&quot;&gt;Lecture Notes&lt;/a&gt;, &lt;a href=&quot;https://www.youtube.com/playlist?list=PLE6Wd9FR--EdyJ5lbFl8UuGjecvVw66F6&quot;&gt;Videos&lt;/a&gt;) by Prof. Nando de Freitas&lt;/li&gt;

&lt;li&gt;Machine Learning Summer School 2013 (Lecture Notes, &lt;a href=&quot;https://www.youtube.com/playlist?list=PLqJm7Rc5-EXFv6RXaPZzzlzo93Hl0v91E&quot;&gt;Videos&lt;/a&gt;)&lt;/li&gt;

&lt;/ol&gt;

&lt;p&gt;&lt;b&gt;Deep Learning&lt;/b&gt;&lt;/p&gt;

&lt;ol&gt;

&lt;li&gt;Deep Learning with PyTorch (&lt;a href=&quot;https://atcold.github.io/pytorch-Deep-Learning/&quot;&gt;Lecture Notes&lt;/a&gt;, &lt;a href=&quot;https://www.youtube.com/playlist?list=PLLHTzKZzVU9eaEyErdV26ikyolxOsz6mq&quot;&gt;Videos&lt;/a&gt;) by Prof. Yann LeCun and Alfredo Canziani&lt;/li&gt;

&lt;li&gt;MIT 6. S191 Introduction to Deep Learning (&lt;a href=&quot;https://introtodeeplearning.com/&quot;&gt;Lecture Notes&lt;/a&gt;, &lt;a href=&quot;https://www.youtube.com/playlist?list=PLtBw6njQRU-rwp5__7C0oIVt26ZgjG9NI&quot;&gt;Videos&lt;/a&gt;)&lt;/li&gt;

&lt;li&gt;Deep Learning (&lt;a href=&quot;https://www.cs.ox.ac.uk/people/nando.defreitas/machinelearning/&quot;&gt;Lecture Notes&lt;/a&gt;, &lt;a href=&quot;https://www.youtube.com/playlist?list=PLE6Wd9FR--EfW8dtjAuPoTuPcqmOV53Fu&quot;&gt;Videos&lt;/a&gt;) by Prof. Nando de Freitas&lt;/li&gt;

&lt;li&gt;DeepMind x UCL 2020 (&lt;a href=&quot;https://www.youtube.com/playlist?list=PLqYmG7hTraZCDxZ44o4p3N5Anz3lLRVZF&quot;&gt;Lectures and Videos&lt;/a&gt;)&lt;/li&gt;

&lt;li&gt;EE-559 – Deep Learning 2019 (&lt;a href=&quot;https://fleuret.org/ee559/&quot;&gt;Lecture Notes and Videos&lt;/a&gt;) by Prof. Fran&amp;ccedil;ois Fleuret&lt;/li&gt;

&lt;/ol&gt;

&lt;p&gt;&lt;b&gt;Reinforcement Learning&lt;/b&gt;&lt;/p&gt;

&lt;ol&gt;

&lt;li&gt; 2021 DeepMind x UCL Reinforcement Learning Lecture Series (&lt;a href=&quot;https://www.youtube.com/playlist?list=PLqYmG7hTraZDVH599EItlEWsUOsJbAodm&quot;&gt;Videos&lt;/a&gt;, Slides below YouTube Video)&lt;/li&gt;

&lt;/ol&gt;

&lt;p&gt;&lt;b&gt;Meta Learning&lt;/b&gt;&lt;/p&gt;

&lt;ol&gt;

&lt;li&gt;CS 330: Deep Multi-Task and Meta Learning 2020 (&lt;a href=&quot;https://cs330.stanford.edu/&quot;&gt;Lectures&lt;/a&gt;, &lt;a href=&quot;https://www.youtube.com/playlist?list=PLoROMvodv4rMC6zfYmnD7UG3LVvwaITY5&quot;&gt;Videos&lt;/a&gt;) by Prof. Chelsea Finn&lt;/li&gt;

&lt;/ol&gt;

&lt;p&gt;&lt;b&gt;Causal Inference&lt;/b&gt;&lt;/p&gt;

&lt;ol&gt;

&lt;li&gt;Introduction to Causal Inference 2020 (&lt;a href=&quot;https://www.bradyneal.com/causal-inference-course&quot;&gt;Lectures and Videos&lt;/a&gt;) by Brady Neal&lt;/li&gt;

&lt;li&gt;Causal inference meets probabilistic models (&lt;a href=&quot;https://www.youtube.com/playlist?list=PLZ_xn3EIbxZEPmFCCCACWe9jpSN6KHA2P&quot;&gt;Videos&lt;/a&gt;)&lt;/li&gt;

&lt;/ol&gt;

&lt;p&gt;&lt;b&gt;Gaussian Processes&lt;/b&gt;&lt;/p&gt;

&lt;ol&gt;

&lt;li&gt;Gaussian Process and Uncertainty Quantification Summer School, 2020 (&lt;a href=&quot;https://gpss.cc/gpss20/program&quot;&gt;Lectures&lt;/a&gt;, &lt;a href=&quot;https://www.youtube.com/playlist?list=PLZ_xn3EIbxZHynuWRdYp4WDtpKm5Xo9Ge&quot;&gt;Videos&lt;/a&gt;)&lt;/li&gt;

&lt;/ol&gt;
</description>
        <pubDate>Fri, 18 Sep 2020 08:16:00 +0000</pubDate>
        <link>https://amootoovaloo.github.io/blog/2020/09/Machine-Learning-Resources</link>
        <guid isPermaLink="true">https://amootoovaloo.github.io/blog/2020/09/Machine-Learning-Resources</guid>
        
        
        <category>Machine Learning and Statistics</category>
        
      </item>
    
      <item>
        <title>Virtual Workshops in 2020</title>
        <description>&lt;p align=&quot;justify&quot;&gt;In 2020, the COVID-19 pandemic moved most conferences and schools online, and working from home became the norm. The virtual format made it possible to attend more events than would otherwise have been practical. In this post, I summarise the virtual workshops and schools I attended that year.&lt;/p&gt;
&lt;ol&gt;
  &lt;li&gt;&lt;u&gt;ICLR 2020&lt;/u&gt; (26&lt;sup&gt;th&lt;/sup&gt; April to 1&lt;sup&gt;st&lt;/sup&gt; May, &lt;a href=&quot;https://iclr.cc/virtual_2020/index.html&quot;&gt;&lt;i style=&quot;font-size:12px&quot; class=&quot;fa&quot;&gt;&amp;#xf08e;&lt;/i&gt;&lt;/a&gt;)&lt;/li&gt;
  &lt;p align=&quot;justify&quot;&gt;I had the opportunity to present my work on weak lensing, data compression and Gaussian Processes (see &lt;a href=&quot;https://slideslive.com/38926202/gaussian-processes-emulator-and-moped-for-weak-lensing&quot;&gt;here&lt;/a&gt;) at ICLR 2020. Beyond my own work, it was inspiring to see how machine learning has grown from a purely scientific field into an engineering discipline that now spans almost every branch of science and engineering. One of the main themes was climate change, one of the most pressing challenges facing society.&lt;/p&gt;
  &lt;li&gt;&lt;u&gt;MLSS 2020&lt;/u&gt; (28&lt;sup&gt;th&lt;/sup&gt; June to 10&lt;sup&gt;th&lt;/sup&gt; July)&lt;/li&gt;
  &lt;p align=&quot;justify&quot;&gt;Although I did not register for this school, all talks were streamed live on YouTube. As my work is closely related to Bayesian analysis, I found the talks by &lt;a href=&quot;https://shakirm.com/&quot;&gt;Shakir Mohamed&lt;/a&gt; particularly valuable (Bayesian Inference I and II). Another notable talk was on meta-learning by &lt;a href=&quot;https://www.stats.ox.ac.uk/~teh/&quot;&gt;Prof. Yee Whye Teh&lt;/a&gt;, in which he quoted the following:&lt;/p&gt;
  &lt;blockquote&gt;
    &lt;p align=&quot;justify&quot;&gt;&lt;small&gt;&lt;i&gt;&quot;Our training procedure is based on a simple machine learning principle: test and train
      conditions must match&quot;&lt;/i&gt; -
      &lt;a href=&quot;https://papers.nips.cc/paper/6385-matching-networks-for-one-shot-learning&quot;&gt;Vinyals et al. 2016&lt;/a&gt;&lt;/small&gt;
    &lt;/p&gt;
  &lt;/blockquote&gt;
  &lt;li&gt;&lt;u&gt;ICML 2020&lt;/u&gt; (12&lt;sup&gt;th&lt;/sup&gt; July to 18&lt;sup&gt;th&lt;/sup&gt; July, &lt;a href=&quot;https://icml.cc/virtual/2020&quot;&gt;&lt;i style=&quot;font-size:12px&quot; class=&quot;fa&quot;&gt;&amp;#xf08e;&lt;/i&gt;&lt;/a&gt;)&lt;/li&gt;
  &lt;p align=&quot;justify&quot;&gt;ICML 2020 offered a wealth of talks, workshops and tutorials. I focused in particular on the Invertible Neural Networks, Normalizing Flows, and Explicit Likelihood Models workshop (&lt;a href=&quot;https://icml.cc/virtual/2020/workshop/5742&quot;&gt;&lt;i style=&quot;font-size:12px&quot; class=&quot;fa&quot;&gt;&amp;#xf08e;&lt;/i&gt;&lt;/a&gt;), which included an interesting talk by &lt;a href=&quot;https://en.wikipedia.org/wiki/Kyle_Cranmer&quot;&gt;Kyle Cranmer&lt;/a&gt; on how deep learning techniques were being applied in science. I also followed the excellent tutorial on Bayesian Deep Learning and a Probabilistic Perspective of Model Construction (&lt;a href=&quot;https://icml.cc/virtual/2020/tutorial/5750&quot;&gt;&lt;i style=&quot;font-size:12px&quot; class=&quot;fa&quot;&gt;&amp;#xf08e;&lt;/i&gt;&lt;/a&gt;) by &lt;a href=&quot;https://cims.nyu.edu/~andrewgw/&quot;&gt;Andrew Wilson&lt;/a&gt;. The conference also offered mentoring sessions on career advice, emerging topics in machine learning, equality and more, which were particularly helpful for a PhD student.&lt;/p&gt;
  &lt;li&gt;&lt;u&gt;OxML 2020&lt;/u&gt; (17&lt;sup&gt;th&lt;/sup&gt; August to 25&lt;sup&gt;th&lt;/sup&gt; August, &lt;a href=&quot;https://www.oxfordml.school/&quot;&gt;&lt;i style=&quot;font-size:12px&quot; class=&quot;fa&quot;&gt;&amp;#xf08e;&lt;/i&gt;&lt;/a&gt;)&lt;/li&gt;
  &lt;p align=&quot;justify&quot;&gt;The school opened with two lectures closely related to my own research: Bayesian Machine Learning by Cheng Zhang and Gaussian Processes by James Hensman. These were followed by lectures on neural networks, natural language processing (NLP), computer vision, representation learning, causal machine learning, reinforcement learning and more. Alongside the lectures, there were tutorials and unconference sessions in which participants took the initiative to discuss specific topics in depth.&lt;/p&gt; 
  &lt;li&gt;&lt;u&gt;GPSS 2020&lt;/u&gt; (14&lt;sup&gt;th&lt;/sup&gt; September to 17&lt;sup&gt;th&lt;/sup&gt; September, &lt;a href=&quot;https://gpss.cc/gpss20/&quot;&gt;&lt;i style=&quot;font-size:12px&quot; class=&quot;fa&quot;&gt;&amp;#xf08e;&lt;/i&gt;&lt;/a&gt;)&lt;/li&gt;
  &lt;p align=&quot;justify&quot;&gt;The school featured lectures on many aspects of Gaussian Processes (GPs), such as scalable GPs and deep GPs, as well as related topics including Bayesian optimisation, kernel design, Bayesian neural networks and composite GPs. &lt;a href=&quot;https://carlhenrik.com/&quot;&gt;Carl Henrik Ek&lt;/a&gt; gave a clear and engaging introduction to GPs.&lt;/p&gt;

&lt;!--   &lt;li&gt;&lt;u&gt;NeurIPS 2020&lt;/u&gt; (6&lt;sup&gt;th&lt;/sup&gt; December to 12&lt;sup&gt;th&lt;/sup&gt; December, &lt;a href=&quot;https://nips.cc/Conferences/2020&quot;&gt;&lt;i style=&quot;font-size:12px&quot; class=&quot;fa&quot;&gt;&amp;#xf08e;&lt;/i&gt;&lt;/a&gt;)&lt;/li&gt;
 --&gt;
&lt;/ol&gt;

&lt;!-- 
&lt;p align=&quot;justify&quot;&gt;This would not have been possible if I were to be physically present, simply because the latter cost at least £200 while with the virtual workshops, it costs no more than £50. I understand that it is not the same as being part of an actual workshop but I personally think I have been able to make the most out of them.&lt;/p&gt; --&gt;

</description>
        <pubDate>Fri, 10 Jul 2020 22:04:00 +0000</pubDate>
        <link>https://amootoovaloo.github.io/blog/2020/07/Virtual-Workshops</link>
        <guid isPermaLink="true">https://amootoovaloo.github.io/blog/2020/07/Virtual-Workshops</guid>
        
        
        <category>Machine Learning and Statistics</category>
        
      </item>
    
      <item>
        <title>How I Take Research to Publication</title>
        <description>&lt;p align=&quot;justify&quot;&gt;In this post, I outline the steps I have found most useful in carrying out research and taking it through to publication. The first first-author paper is generally the most challenging. It was through this process that I learned to think critically and to answer questions such as &apos;Why is this research important?&apos;, &apos;Where will it be used?&apos;, &apos;How does it differ from previous work?&apos;, &apos;What are its new contributions to the field?&apos; and &apos;How do I manage my time effectively?&apos;&lt;/p&gt;

&lt;ol&gt;
	&lt;li&gt;&lt;b&gt;Project definition&lt;/b&gt;&lt;/li&gt;
	&lt;p align=&quot;justify&quot;&gt;The first step is to define the project. There is, however, always an element of risk at this stage: researchers sometimes find that their projects no longer seemed promising after several months of work.&lt;/p&gt;	
	&lt;li&gt;&lt;b&gt;Brainstorming&lt;/b&gt;&lt;/li&gt;
	&lt;p align=&quot;justify&quot;&gt;To avoid the pitfalls described above, I strongly recommend holding brainstorming sessions with your collaborators to identify possible risks and limitations.&lt;/p&gt;
	&lt;blockquote&gt;
	&lt;p align=&quot;justify&quot;&gt;&lt;small&gt;&lt;i&gt;&quot;If I had an hour to solve a problem I’d spend 55 minutes thinking about the problem and five minutes thinking about solutions.&quot;&lt;/i&gt; - Albert Einstein&lt;/small&gt;&lt;/p&gt;
	&lt;/blockquote&gt;
	&lt;li&gt;&lt;b&gt;Identify a main reference paper&lt;/b&gt;&lt;/li&gt;
	&lt;p align=&quot;justify&quot;&gt;Finding a main reference paper is fundamental to the whole process. This means understanding the paper and its mathematics, and being able to implement (code) at least part of it.&lt;/p&gt;
	&lt;li&gt;&lt;b&gt;Fill in the gaps&lt;/b&gt;&lt;/li&gt;
	&lt;p align=&quot;justify&quot;&gt;At this point, we typically have only a blurred picture of how to develop the main idea, largely because of gaps in our knowledge of specific topics. It is very helpful to master these topics (through textbooks, lecture notes, online resources, YouTube and GitHub) to fill in the gaps.&lt;/p&gt;
	&lt;li&gt;&lt;b&gt;Experiments&lt;/b&gt;&lt;/li&gt;
	&lt;p align=&quot;justify&quot;&gt;This is probably the most important part of the entire process, where we write and test our code. It can seem daunting, as we will inevitably encounter many failures. However,&lt;/p&gt;	
	&lt;blockquote&gt;
	&lt;p align=&quot;justify&quot;&gt;&lt;small&gt;&lt;i&gt;&quot;Failure provides the opportunity to begin again, more intelligently.&quot;&lt;/i&gt; - Henry Ford&lt;/small&gt;&lt;/p&gt;
	&lt;/blockquote&gt;
	&lt;li&gt;&lt;b&gt;Updates&lt;/b&gt;&lt;/li&gt;
	&lt;p align=&quot;justify&quot;&gt;Crucially, it is good practice to document our code and draft the paper as we go. This allows us to keep track of everything we have done since the start of the project.&lt;/p&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;b&gt;The Paper&lt;/b&gt;&lt;/p&gt;
&lt;p align=&quot;justify&quot;&gt;Different journals require different formats, so it is important to identify and adopt the right format from the very beginning to work efficiently. Below are some brief suggestions I received from my supervisors:&lt;/p&gt;
&lt;ol&gt;
	&lt;li&gt;&lt;u&gt;Abstract&lt;/u&gt;: Summarise key findings. Numbers are essential.&lt;/li&gt;
	&lt;li&gt;&lt;u&gt;Introduction&lt;/u&gt;: Motivate the study.&lt;/li&gt;
	&lt;li&gt;&lt;u&gt;Body&lt;/u&gt;: Elaborate on the theory, data, methods and results.&lt;/li&gt;
	&lt;li&gt;&lt;u&gt;Conclusion&lt;/u&gt;: Remind the reader what the paper is about and briefly explain how its aims have been met.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;b&gt;Most Important Suggestions&lt;/b&gt;&lt;/p&gt;

&lt;ol&gt;
	&lt;li&gt;Keep in regular contact with your supervisors (emails and weekly meetings).&lt;/li&gt;
	&lt;li&gt;Have at least one mentor. I have found it easy to talk to my collaborator, who is a Research Fellow.&lt;/li&gt;
	&lt;li&gt;Always take notes. We live in a world with a wealth of information, and it is important to organise ideas and information.&lt;/li&gt;
	&lt;li&gt;Do not work in isolation. In other words, do not pigeonhole yourself.&lt;/li&gt;
	&lt;li&gt;Ask questions early and often; resolving a misunderstanding quickly saves time later.&lt;/li&gt;
	&lt;li&gt;Understanding is key. Some work is largely engineering, but we should still be able to explain the concepts behind it.&lt;/li&gt;
	&lt;li&gt;Acknowledge when you do not understand something; when we seek knowledge, help and support, we will receive it.&lt;/li&gt;
&lt;/ol&gt;
</description>
        <pubDate>Fri, 10 Jul 2020 18:08:00 +0000</pubDate>
        <link>https://amootoovaloo.github.io/blog/2020/07/Research-Paper</link>
        <guid isPermaLink="true">https://amootoovaloo.github.io/blog/2020/07/Research-Paper</guid>
        
        
        <category>Machine Learning and Statistics</category>
        
      </item>
    
  </channel>
</rss>
