Enquire Now
70+ Topics · Spectre · Spectre · cloud sim Sim · MATLAB · Webots · Hardware · Bangalore 2026

Third Harmonic Injection Pwm Matlab

Simulation · Control · Perception · Hardware — 12 Lead ECG Acquisition — hardware, sensors, cloud dashboards and protocols (Spectre, REST, CoAP, WebSockets) for BE BTech MTech students. Final-year robotics support with Spectre stacks, simulation worlds, reports and viva from Bangalore.

70+
Related Topics
6+
Sim & HW Tools
4.9★
573 Ratings

Abstract

Reinforcement Learning (RL) has made significant strides in complex tasks but struggles in multi-task settings with different embodiments. World model methods offer scalability by learning a simulation of the environment but often rely on inefficient gradient-free optimization methods for policy extraction. In contrast, gradient-based methods exhibit lower variance but fail to handle discontinuities.

third-harmonic-injection-pwm-matlab Diagram
Figure: System Model & Simulation Flow for Third Harmonic Injection Pwm Matlab

Our work reveals that well-regularized world models can generate smoother opti- mization landscapes than the actual dynamics, facilitating more effective first-order optimization. We introduce Policy learning with multi-task World Models (PWM), a novel model-based RL algorithm for continuous control. Initially, the world model is pre-trained on offline data, and then policies are extracted from it using first-order optimization in less than 10 minutes per task. PWM effectively solves tasks with up to 152 action dimensions and outperforms methods that use ground- truth dynamics. Additionally, PWM scales to an 80-task setting, achieving up to 27% higher rewards than existing baselines without relying on costly online planning. Visualizations and code are available at imgeorgiev.com/pwm.

third-harmonic-injection-pwm-matlab Diagram
Figure: System Model & Simulation Flow for Third Harmonic Injection Pwm Matlab

Pwm (Ours)

Figure 1: We propose PWM, a new method for multi-task RL that utilizes pre-trained world models to learn policies for each task. When sufficiently regularized, these world models induce smooth optimization landscapes, which allows for efficient first-order optimization. Our approach can solve tasks in <10 minutes and achieves higher rewards in both single-task and multi-task environments.

third-harmonic-injection-pwm-matlab Diagram
Figure: System Model & Simulation Flow for Third Harmonic Injection Pwm Matlab

Introduction

The pursuit of generalizability in machine learning has recently been propelled by the training of large models on substantial datasets Brown et al. (2020); Kirillov et al. (2023); Bommasani et al. (2021). Such advancements have notably permeated robotics, where multi-task behavior cloning techniques have shown remarkable performance Zitkovich et al. (2023); Octo Model Team et al.

third-harmonic-injection-pwm-matlab Diagram
Figure: System Model & Simulation Flow for Third Harmonic Injection Pwm Matlab

(2024); Goyal et al. (2023); Bousmalis et al. (2023). Nevertheless, these approaches predominantly hinge on near-expert data and struggle with adaptability across diverse robot morphologies due to their dependence on teleoperation Zitkovich et al. (2023); Octo Model Team et al. (2024); Kumar et al. (2021).

third-harmonic-injection-pwm-matlab Diagram
Figure: System Model & Simulation Flow for Third Harmonic Injection Pwm Matlab

In contrast, Reinforcement Learning (RL) offers a robust framework capable of learning from suboptimal data, addressing the aforementioned limitations. However, traditional RL has been focused on single-task experts Mnih et al. (2013); Schulman et al. (2017); Haarnoja et al. (2018).

third-harmonic-injection-pwm-matlab Diagram
Figure: System Model & Simulation Flow for Third Harmonic Injection Pwm Matlab

Recently, Hansen et al. (2024) suggested that a potential pathway to multi-task RL is with the world

Published As A Conference Paper At Iclr 2025

models framework, where a large model learns the environment dynamics and is then combined with Zeroth-order Gradient (ZoG) methods. Despite advancements, ZoG methods struggle with sample inefficiency due to the high variance Mohamed et al. (2020); Suh et al. (2022); Parmas et al. (2023) and online planning time scales with model size, rendering it infeasible at scale.

third-harmonic-injection-pwm-matlab Diagram
Figure: System Model & Simulation Flow for Third Harmonic Injection Pwm Matlab

First-order Gradient (FoG) methods provide a low-variance alternative that have shown superior sample efficiency and asymptotic performance when combined with smooth differentiable simulations Xu et al. (2022). However, they struggle to optimize through discontinuities Suh et al. (2022); Georgiev et al. (2024). In this work, we explore the tight coupling between FoG optimization and world models through the lens of differentiable simulation. Counter-intuitively, we find that for gradient-based optimization, we don’t want world models to be accurate; instead, we want them to be smooth and have a low optimality gap. This in turn enables efficient FoG optimization.

third-harmonic-injection-pwm-matlab Diagram
Figure: System Model & Simulation Flow for Third Harmonic Injection Pwm Matlab

Building on these insights, we propose Policy learning with multi-task World Models (PWM), an algorithm that can learn policies from offline pre-trained world models in under <10 minutes per task. With this new-found efficiency, we also propose a new multi-task framework, where instead of training a multi-task algorithm, we only train a multi-task world model and then extract a policy for each task. This decoupling of the supervised and RL objectives results in more stable and efficient learning with higher episode rewards. Our empirical evaluations on high-dim. tasks indicate that PWM not only achieves higher reward than baselines but also outperforms methods that use ground-truth dynamics. In a multi-task scenario utilizing a pre-trained 48M parameter world model from TD-MPC2, PWM achieves up to 27% higher reward than TD-MPC2 without relying on online planning. This underscores the efficacy of PWM and supports our broader contributions: 1. Correlation Between World Model Smoothness and Policy Performance: Through pedagogical examples and ablations, we demonstrate that smoother, better-regularized world models significantly enhance policy performance. Notably, this results in an inverse correlation between model accuracy and policy performance.

third-harmonic-injection-pwm-matlab Diagram
Figure: System Model & Simulation Flow for Third Harmonic Injection Pwm Matlab

2. Efficiency of First-Order Gradient (FoG) Optimization: We show that combining FoG optimization with well-regularized world models enables more efficient policy learning compared to zeroth-order methods. Furthermore, policies learned from world models asymp- totically outperform those trained with ground-truth simulation dynamics, emphasizing the importance of the tight relationship between FoG optimization and world model design.

third-harmonic-injection-pwm-matlab Diagram
Figure: System Model & Simulation Flow for Third Harmonic Injection Pwm Matlab

3. Scalable Multi-Task Algorithm: Instead of training a single multi-task policy model, we propose PWM, a framework where a multi-task world model is first pre-trained on offline data. Then per-task expert policies are extracted in <10 minutes per task, offering a clear and scalable alternative to existing methods focused on unified multi-task models.

third-harmonic-injection-pwm-matlab Diagram
Figure: System Model & Simulation Flow for Third Harmonic Injection Pwm Matlab

Background

We focus on discrete-time and infinite-horizon Reinforcement Learning (RL) scenarios characterized by system states s ∈Rn = S, actions a ∈Rm = A, dynamics function f : S ×A →S and a reward function r : S × A →R. Combined, these form a Markov Decision Problem (MDP) summarized by the tuple (S, A, f, r, γ) where γ is the discount factor. Actions at each timestep t are sampled from a stochastic policy at ∼πθ(·|st), parameterized by θ. The goal of the policy is to maximize the

(1)

where ρ(s1) is the initial state distribution. Since this maximization over an infinite sum is intractable, in practice we often maximize over a value estimate. The value of a state st is defined as the expected

(2)

When V is approximated with a learned model with parameters ψ and πθ attempts to maximize some function of V , we arrive at the popular and successful actor-critic architecture Konda & Tsitsiklis (1999). Additionally, in MBRL it is common to also learn approximations of f and r, which we denote as Fϕ and Rϕ, respectively. It has also been shown to be beneficial to encode the true state

𝜃

(a) Ball-wall visualization.

Simnorm

(b) Problem landscape.

Error

Opt.

3.47

(c) Model error and optimality gap. Figure 2: Ball-wall pedagogical example. The left figure visualizes the problem. The middle figure shows the problem landscape induced by each model. J(θ) shows the true underlying function and the two other are MLPs with different activation functions. We minimize each of these problems using gradient descent and starting at θ = −π (marker ×). The colored crosses represent the solutions converged to for each model. The right table shows the model approximation error during training and the optimality gap |J(θ∗) −J(ˆθ)| between the global minimum θ∗and the solution found for each model ˆθ.

third-harmonic-injection-pwm-matlab Diagram
Figure: System Model & Simulation Flow for Third Harmonic Injection Pwm Matlab

s into a latent state z using a learned encoder Eϕ Hafner et al. (2019); Hansen et al. (2022; 2024); Hafner et al. (2023). Putting together all of these components, we can define a model-based actor- critic algorithm to consist of the tuple (πθ, Vψ, Eϕ, Fϕ, Rϕ) which can describe popular approaches such as Dreamer Hafner et al. (2019; 2023) and TD-MPC2 Hansen et al. (2024). Notably, we make an important distinction between the types of components. We refer to Eϕ, Fϕ and Rϕ as the world model components since they are supervised learning problems with fixed targets. On the other hand, πθ and Vψ optimize for moving targets, which is fundamentally more challenging, and we refer to them as the policy components.

third-harmonic-injection-pwm-matlab Diagram
Figure: System Model & Simulation Flow for Third Harmonic Injection Pwm Matlab

Policy Optimization Through World Models

This paper builds on the insight that since access to Fϕ and Rϕ is assumed through a pre-trained world model, we have the option to optimize Eq. 1 via First-order Gradient (FoG) optimization which exhibit lower gradient variance, more optimal solutions, and improved sample efficiency Mohamed et al. (2020). In our setting, these types of gradients are obtained by directly differentiating the expected terms of Eq. 1 as shown in Eq. 3. Note that this gradient estimator is also known as reparameterized gradient Kingma et al. (2015) and pathwise derivative Schulman et al. (2015). While we use the explicit ∇ notation below, we later drop it for simplicity as all gradient types in this work are first-order gradients.

third-harmonic-injection-pwm-matlab Diagram
Figure: System Model & Simulation Flow for Third Harmonic Injection Pwm Matlab

As ∇

θ J(θ) in itself is a random variable, we need to estimate it. A popular way to do that in practice is via Monte-Carlo approximation, where we are interested in two properties: bias and variance. In Sections 3.1 and 3.2 we tackle each aspect with a toy robotic control problem to build intuition. In Section 3.3, we combine our findings to propose a new algorithm.

third-harmonic-injection-pwm-matlab Diagram
Figure: System Model & Simulation Flow for Third Harmonic Injection Pwm Matlab

= ∇J(Θ), Only If Both The Dynamics F And

rewards r are Lipschitz-smooth Suh et al. (2022). However, many robotic problems involving contact are inherently non-smooth, which breaks these conditions and results in gradient sample error where

̸

= ∇J(θ) under a finite number of samples N. Instead of directly optimizing the true, discontinuous objective, it is advantageous to optimize a smooth surrogate, such as a model learned by a regularized deep neural network.

To illustrate this concept, we use a toy problem where a ball is thrown toward a wall at a fixed velocity, as shown in Figure 2a. The objective is to find the optimal initial angle θ such that we maximize

Mlp H=16

Figure 3: Double pendulum pedagogical example. The middle figure evaluates the variance of policy gradient estimates over N = 100 Monte-Carlo samples for varying horizons H. The right figure shows the same data but plots the Expected Signal-to-Noise ratio (ESNR) with higher values translating to more useful gradients. These results suggests that world models trained over long horizon trajectories provide more useful gradients. Note that H = 3 and H = 16 in the figure legends refer to the training horizon of the models.

forward distance. In this simplified pedagogical example, we assume that the ball "sticks" to the wall, creating a discontinuous optimization landscape (Figure 2b). We compare the performance of two models in approximating this objective: a 2-layer Multi-Layer Perceptron (MLP) with ReLU activation and another MLP with SimNorm activation Hansen et al. (2024) in the intermediate layers.

SimNorm normalizes a latent vector z by projecting it into simplices with dimension P using a softmax operator. Given an input vector z, SimNorm can be expressed as a mapping into L vectors:

(4)

We trained both MLPs and observed their effects on smoothing the optimization landscape (Figure 2b). The ReLU-activated MLP smooths the landscape but introduces a local minimum that hinders gradient descent, particularly when starting from θ = −π, resulting in a large optimality gap (difference between the solution and the optimal solution: ∥ˆθ −θ∗∥). In contrast, the SimNorm-activated MLP has additional regularization, which reduces the optimality gap at the expense of model accuracy (Table 2c). This example highlights that more accurate models do not always lead to better policies, as noted by Lambert et al. (2020). Our findings extend this by showing that for FoG optimization, prioritizing smoothness over accuracy can lead to improved results. Further details are provided in Appendix A.

Learning With Chaotic Dynamics

While FoGs have lower variance per step, they can accumulate significant variance over long-horizon rollouts Metz et al. (2021). Suh et al. (2022) link this variance to the smoothness of models and

∝∥∇F(S, A)∥2H. At Sufficiently High H, The

high variance renders FoGs ineffective in chaotic systems. Chaotic systems are characterized by their sensitivity to initial conditions, where small perturbations can lead to exponentially divergent trajectories, making long-term prediction particularly challenging. The double pendulum, also known as the Acrobot Murray & Hauser (1991), is a classic example of such a system (Figure 3).

We analyze the variance of gradient estimators in the double pendulum using both the true dynamics and a SimNorm-activated MLP model. The MLP model was trained for auto-regressive prediction horizons of H = 3 and H = 16 until convergence on a large dataset. Figure 3 shows that both learned models exhibit reduced variance compared to the true dynamics. However, as noted by Parmas et al. (2023), variance alone is insufficient for drawing definitive conclusions about gradient quality. Instead, they propose analyzing gradients via their Expected Signal-to-Noise Ratio (ESNR),

(5)

In Figure 3, we observe that learned models exhibit higher ESNR than the true dynamics, providing more useful gradients. Notably, the training horizon plays a critical role, with the H = 16 model sustaining a higher ESNR over higher H. We conclude that learned world models offer more informative policy gradients than the true system dynamics. Further details in Appendix B.

Φ ←Φ −Αϕ∇Lwm(Φ)

▷Eq.

Θ ←Θ + Αθ∇Lπ(Θ)

▷Eq.

Ψ ←Ψ −Αψ∇Lv (Ψ)

▷Eq.

Differentiable Physics Simulators That Provide

gradients with low sample error and variance.

Via Ddpg-Style Gradients:

∇θJ(θ) ≈Ea∼π(·|s)[∇θQ(s, a)].

Ficiently Learning Policies From Large Multi-Task

world models. Framework.

From Multiple Tasks, We First Train A Multi-Task

world model to predict future states and rewards.

Then For Each Task We Want To Solve, We Learn A

single policy in minutes using FoG optimization.

The Policy Is Then Deployed To Solve The Task And

optionally fine-tune its world model and policy.

Policy Actor-Critic Approach Inspired By Dif-

ferentiable simulation approaches Xu et al.

Propagated Through The World Model, While The

critic is trained via TD(λ). The key to our approach is that training is done in a batched fashion where multiple trajectories are imagined in parallel. The actor loss function is akin to Eq. 1 but features rewards over a fixed horizon H, terminal value bootstrapping and usage of the learned world model

(6)

The critic is trained in a model-free fashion using TD(λ) over an H-step rollout in latent space z as seen in other similar on-policy methods Sutton & Barto (2018); Hafner et al. (2019); Xu et al. (2022):

(9)

We use an ensemble of 3 critics to reduce variance. To enable FoG optimization, it is important to use

Model Proposed By

TD-MPC2 Hansen et al. (2024) with learnable task embeddings e. It is trained in an auto-regressive fashion by sampling data from a buffer with loss function:

(10)

where sg(·) is the stop-gradient operator and CE is the cross-entropy loss function. Reward prediction is formulated as a discrete regression problem in log-transformed space. Furthermore, Eϕ and Fϕ use SimNorm activation (Eq. 4) in their output layers. All trainable models are fully-connected MLPs with LayerNorm Ba et al. (2016) and Mish activation Misra (2019). The complete algorithm is shown in Algorithm 1. Further implementation details can be found in Appendix C.

Published As A Conference Paper At Iclr 2025

Figure 4: High-dimensional single-task environments (left to right): Hopper, Ant, Anymal, Humanoid and SNU Humanoid. Our method successfully learns tasks with up to m = 152 continuous action dimensions. Additional 80 multi-task environments used in this paper are listed in Appendix E

Contact-Rich Single Tasks

The aim of this section is to understand whether smooth world models create better optimization landscapes than ground-truth dynamics, facilitating efficient FoG optimization. We study this on 5 continuous control tasks (Figure 4) with up to A = R152 using the differentiable simulator dflex Xu et al. (2022). Comparisons include SHAC Xu et al. (2022), which uses ground-truth dynamics and rewards with a similar actor-critic architecture to PWM. Furthermore, we compare against two world model approaches. DreamerV3 Hafner et al. (2023) learns its world model via reconstruction, its actor via ZoG optimization, and critic via Model-based Value Expansion (MVE) Feinberg et al.

(2018). TD-MPC2 Hansen et al. (2024) uses the same world model as PWM but learns a policy in a model-free fashion and actively plans at inference time. We additionally include model-free baselines PPO Schulman et al. (2017) and SAC Haarnoja et al. This comparison allows us to understand whether (1) FoG-based optimization can learn better policies asymptotically and (2) whether smooth world models induce better optimization landscapes for FoG optimization.

We conduct this experiment across 5 tasks with 10 seeds each, where PWM, DreamerV3, and TD- MPC2 use pre-trained world models and are left to learn a policy and fine-tune their world models online. This is done to enable a fair comparison to SHAC, which directly has access to the simulation model and does not require any training. The results in Figure 5 reveal that (1) PWM is able to learn policies with higher reward than SHAC asymptotically, indicating that regularized world models induce smoother optimization landscapes than the true (discontinuous) dynamics. Furthermore (2), despite using the same world model, our method is able to learn policies with higher rewards than TD-MPC2 without the need for online planning. However, PWM does not scale well to the highest-dimensional task. More experiment details and results are included in Appendix D.

Pwm (Ours)

Figure 5: Aggregate results from high-dimensional locomotion tasks where each agent is trained to solve just that task (i.e. specialist). The left figure summarizes rewards achieved at the end of training using 50% IQM for the solid lines and 95% CI as suggested by Agarwal et al. (2021), as well as mean for the dashed lines. We see that PWM achieves higher rewards than our main baselines TD-MPC2 and SHAC. The right figure shows score distributions across all tasks which lets us understand the performance variability of each approach.

Pwm (Ours)

Figure 6: Multi-task results. The left figure shows results of multi-task agents in the 30 and 80 task set settings which include environments from dm_control Tunyasuvunakool et al. (2020) and MetaWorld Yu et al. (2020). The results show 50% IQM with the solid lines and mean with the dashed lines. The bars represent 95% CI. In both settings PWM achieves higher reward than TD-MPC2 without the need for online planning. The middle figure compares the training and inference times of TD-MPC2 and PWM for the 48M parameter model. PWM has significantly lower inference time as it does not plan online. The right figure shows a comparison between multi-task PWM and TD-MPC2 and single-task experts SAC and DreamerV3 on the MT30 task set. Notably, PWM is able to match the performance of SAC and DreamerV3.

Multi-Task World-Model

We analyze the scalability of our proposed framework and method to large multi-task pre-trained world models. We evaluate on two settings: (1) 30 continuous control dm_control tasks Tunyasuvunakool et al. (2020) ranging from m = 1 to m = 6 and (2) 80 tasks, which include 50 additional manipulation tasks from MetaWorld Yu et al. (2020) with n = 39 and m = 4. These two multi-task settings were introduced as MT30 and MT80 by Hansen et al. (2024). In conducting our experiments, we harness the same data and world model architecture as TD-MPC2. The data consists of 120k and 40k trajectories per dm_control and MetaWorld task, respectively generated by 3 random seeds of TD-MPC2 runs. The world models we use are the 48M parameter models introduced in Hansen et al.

(2024) with slight modifications to make them differentiable (Appendix C). To train PWM, we first pre-train the world models on the dataset in a similar fashion to TD-MPC2, but with training H = 16 and γ = 0.99 for better first-order gradients, as highlighted in Section 3.2.

Then we train a PWM policy on each particular task using the offline datasets for 10k gradient steps, which take 9.3 minutes on an Nvidia RTX6000 GPU. We evaluate task performance for 10 seeds for each task and aggregate results in Figure 6. We compare against TD-MPC2, which learns a multi-task policy while pre-training its world model and relies on online planning at inference. We can see that PWM learns behavior, achieving a higher reward than TD-MPC2 while also being significantly faster at inference time. While the fast per-task training is enabled by FoG optimization, we also find that training a single multi-task policy produces poor results, as shown in Appendix F. We further compare our multi-task PWM policy to online-trained single-task experts SAC Haarnoja et al. (2018) and DreamerV3 Hafner et al. (2023). Figure 6 reveals that multi-task PWM, while disadvantaged, performs comparably to the single-task experts without requiring any environment interaction and only training policies for ≤10 minutes per task. Additional results in Appendix E.

Ablations

We perform 4 ablations on the complex single task experiments in order to understand the nuances of first-order optimization through world models with PWM. We increase the contact stiffness to be more realistic, but also more stiff contact gives gradients with high sample error Suh et al. (2022). We run the same experiment as Section 4.1, but only for the Hopper task, and present the aggregate results from 5 random seeds in Figure 7a, where we normalize rewards by the maximum reward achieved by PPO in Section 4.1. We see that while PPO and PWM rewards remain similar to prior results, SHAC performance decreases by 48%. This shows that regularized world models are robust to stiff contact models and thus more generalizable than differentiable simulations.

Pwm

(a) Contact stiffness ablation.

2048

(b) Policy batch size ablation. Figure 7: Left figure shows contact stiffness ablation where we increase contact stiffness on the Hopper task and analyze the effects on policy learning. The results indicate that stiff (but realistic) contact has adverse effects on SHAC which uses the simulation model to learn. Meanwhile, PPO and PWM remain unaffected with PWM still obtaining 17% more reward than PPO asymptotically. The right figure shows a policy batch size ablation on the Any task where we vary only the batch size used to train the policy components of PWM. Unfortunately we observe that PWM provides best result within a unit of time by using small batch sizes. Both figures show 50% IQM and 95% CI over 5 random seeds.

World Model Loss

(a) World model ablation.

Pwm

(b) World model vs policy sample efficiency. Figure 8: The left figure ablates the activation functions of the world model used to learn policies on the Ant task. We progressively add more regularization to the world model via changes to the activation function and observe an inverse correlation between world model loss and policy reward. This indicates that we should not construct world models for accuracy but for policy learning. The right figure investigates the policy sample efficiency on 5 dm_control tasks. We use the same data to pre-train world models for varying amount of gradient steps and then train the policy for 50k gradient steps and compare against TD-MPC2 (without planning). The results indicate that PWM policies are significantly more sample efficient but also require better trained world models. All results shown are 50% IQM with 95% CI across 5 random seeds.

The second ablation explores batch sizes for policy learning with first-order gradients. Contrary to model-free methods, which can scale to large batch sizes, we find that FoG techniques like PWM benefit from smaller batch sizes. We explore this on the Ant task in Figure 7b, where we plot 50% IQM rewards over 5 random seeds. While larger batch sizes allow us to generate more data within a unit of time, that does not necessarily translate to learning better policies.

Next we ablate the world model regularization. We perform the same experiment as Section 4.1 on the Ant task but now pre-train 3 different world models. (1) with ReLU activation func., (2) with Mish activation func. and (3) with Mish activation func. and SimNorm activation func. at the output layers of Eϕ and Fϕ. Figure 8a reveals that while less regularization results in lower world model error, that does not translate to learning better policies. Surprisingly, less regularized world models enable policies to start faster (up to 1M steps) but plateau to a suboptimal policy. Additional results in Appendix F.

To understand the policy sample efficiency of PWM while controlling for the world model, we perform an ablation where we pre-train the same world model for [50k, 100k, 250k] gradient steps on an offline dataset. Then we fix the world model and train only the policy components on the

Published As A Conference Paper At Iclr 2025

same dataset for 50k gradient steps and measure the reward. We do this for 3 random seeds and 5 dm_control tasks. We repeat the same experiment for TD-MPC2 but disable its planning component in order to understand the learning dynamics of each method’s policy components. The results in Figure 8b show that the PWM policy components are significantly more sample efficient than TD-MPC2 but also require better trained world models in order to obtain high reward.

Related Work

Reinforcement learning (RL) strategies are divided into model-based and model-free approaches, with the latter not assuming a model of the environment Arulkumaran et al. (2017). Model-free approaches, such as PPO Schulman et al. (2017) and SAC Haarnoja et al. (2018) do not require a model of the environment and represent on-policy and off-policy methods, respectively. These algorithms use an actor-critic structure, where the critic assesses the policy while the actor updates it through gradient-based optimization to maximize rewards Konda & Tsitsiklis (1999).

Gradient estimator types. In the absence of direct access to dynamics and reward functions, it is common to use the Policy Gradient Theorem Sutton et al. (1999), a zeroth-order method, to estimate gradients. Although robust to discontinuities, this method exhibits high variance, leading to sample inefficiency Mohamed et al. (2020). In contrast, first-order gradients (FoG) offer lower variance by differentiating through the objective but struggle with discontinuities Suh et al. (2022). Differentiable simulations have risen as a tool to study the properties of gradient estimators Howell et al. (2022); Metz et al. (2021) and have produced model-based algorithms that use FoG optimization through physics to learn high-performing policies Xu et al. (2022); Georgiev et al. (2024).

Multi-task models. While traditional RL focuses on single-task policies, the broader robotics field is increasingly adopting large multi-task models through behavior cloning Firoozi et al. (2023). Recent efforts like Open X Padalkar et al. (2023) and Octo Octo Model Team et al. (2024) have demonstrated improved performance across various tasks and embodiments by leveraging large models and datasets.

However, the potential of these large-scale approaches in RL remains largely unexplored. While GATO Reed et al. (2022) attempted to scale model-free RL across multiple tasks, it faced challenges with sample inefficiency and required significant fine-tuning. Conversely, TD-MPC2 Hansen et al.

(2024) successfully scaled a 317M parameter world model for online planning across 80 tasks. While showing impressive multi-task scalability, it failed to solve all tasks and exhibits limited scalability due to online planning. Our work builds on the world model architecture proposed by TD-MPC2 but employs FoG optimization for policy learning and extracts per-tasks policies. DreamerV3 Hafner et al. (2023) also integrates world models with FoG but focuses on online learning without addressing multi-task scenarios. Our work delves deeper into the relationship between world models and policy learning, exploring the essential characteristics of world models that facilitate efficient optimization.

Conclusion

In this study, we analyzed world models through policy gradient estimation and identified an inverse correlation between the accuracy of world models and episode rewards. We concluded that world models should prioritize smoothness and a smaller optimality gap over accuracy to enhance policy performance. Building on these insights, we propose Policy learning through Multi-task World Models (PWM), a MBRL algorithm that integrates smooth world models with first-order gradient (FoG) optimization. Our evaluations showed that PWM can outperform existing methods, including those with access to ground-truth simulation dynamics, in learning high-reward policies for high- dimensional tasks. To scale to a multi-task settings, we propose a framework where world models are pre-trained offline and treated as differentiable simulations. Our results demonstrate that PWM can be used to learn expert policies in <10 minutes per task, achieving higher rewards without the need for expensive online planning. With ample data and large, smooth world models, we believe this approach has significant potential for scalability.

Limitations. Despite its demonstrated efficacy, PWM has notable limitations. Firstly, performance relies heavily on the availability of substantial pre-existing data to train the world model, which might not always be feasible, especially in novel or low-data environments. Secondly, although PWM facilitates fast and cost-effective policy training, it necessitates re-training for each new task, which could limit its applicability in scenarios requiring rapid adaptation to diverse tasks. Lastly, the current TD-MPC2 world models used are difficult to train at scale due to their auto-regressive formulation.

Published As A Conference Paper At Iclr 2025

Reproducibility statement. Code, training data and checkpoints are made available at policy-world-model.github.io. We rely on dflex, MetaWorld, DMControl and MuJoCo for simulation which are publicly available under MIT and Apache 2.0 licenses. We use multi-task data from TD-MPC2 which is publicly available. Implementation details and full list of hyper- parameters are available in Appendix C.

Deployment In Out-Of-Position Situations

D. Bendjaballah1, A. Bouchoucha1, M. L. Sahli1,2* and J-C. Gelin2

Abstract

Side-impact collisions represent the second greatest cause of fatality in motor vehicle accidents. Side-impact airbags have been installed in recent model year vehicle due to its effectiveness in reducing passengers’ injuries and fatality rates. In meeting these requirements, simulations of folding and deploying airbags are very useful and are widely used. The paper presents a simulation method for the deploying airbags using three materials in different working conditions. Finite element analysis is primarily used to evaluate this concept. In these simulations, the gas flow is described by the conservation laws of mass, momentum, and energy. The numerical results indicate that the FE method in this paper is capable of capturing airbag deploying process accurately.

ansys-airbag-injury-simulation Diagram
Figure: System Model & Simulation Flow for Ansys Airbag Injury Simulation

Keywords: Airbag simulations, Out-of-position, Crash, Modeling, Out-of-position

Background

The passive safety of cars has become a very high prior- ity issue for the automotive industry. Today, there are not only one or two airbags in a car; certain models have ten times more than that. With the increasing usage of airbags, the number of accidents where the airbag itself can cause an injury to the occupant also increases

(Augenstein Et Al. 2003; Gabauer And Gabler 2010;

Audrey et al. 2011). As is well known, safety belts are also now devices designed to provide protection to the users of vehicles during crash events, minimizing the loads necessary to adapt their movement to the move- ment of the car (Freesmeier and Butler 1999; Schmitt et al. 1997). In general, the seat belt is designed to restrain the occupant in the vehicle and prevent the

Occupant From Having Harsh Contacts With Interior

surfaces of the vehicles. The airbag acts to cushion any impact with vehicle structure and has positive internal pressure, which can exert distributed restraining forces over the head and face. As a safety component of auto- mobile, an airbag decreases occupants’ injury likelihood effectively in case of an accident (Ruff et al. 2007). These safety elements can reduce the death rates on the roads, and its protection effects have been widely approved (Crandall et al. 2001; Teru and Ishikawa 2003). With computational tools such as finite element methods designed for dynamic contact problems, crashworthiness simulations can now be used with reliable accuracy to evaluate occupant protection in various collision condi- tions with safety metric/parameters such as acceleration, head injury criteria, intrusion distance, intrusion vel- ocity, and neck forces (neck injury risk or whiplash).

ansys-airbag-injury-simulation Diagram
Figure: System Model & Simulation Flow for Ansys Airbag Injury Simulation

Thus, new types of airbag products are being developed to handle different collision scenarios.

Become Standard Equipment On Most New Passenger

vehicles (Braver and Kyrychenko 2004; Teng et al. 2007; Yoganandan et al. 2007). The airbag cushion is com- posed of a woven fabric which is rapidly inflated during a car crash. The airbag dissipates the passenger’s kinetic energy thereby reducing injury through biaxial stretching of the fabric bag and escaping gas through vents. There- fore, the performance of the airbag is greatly influenced by the mechanical properties of the fabric. Generally, air bags are designed to deploy in a crash that is equivalent to a vehicle crashing into a solid wall at 8 to 14 mph.

ansys-airbag-injury-simulation Diagram
Figure: System Model & Simulation Flow for Ansys Airbag Injury Simulation

Air bags most often deploy when a vehicle collides with another vehicle or with a solid object like a tree. There are various types of airbags: frontal, side-impact, and curtain airbags. In general, the passenger side airbags are usually larger than the driver airbags (see Fig. 1).

ansys-airbag-injury-simulation Diagram
Figure: System Model & Simulation Flow for Ansys Airbag Injury Simulation

Besançon, France

© The Author(s). 2017 Open Access This article is distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons license, and indicate if changes were made.

ansys-airbag-injury-simulation Diagram
Figure: System Model & Simulation Flow for Ansys Airbag Injury Simulation

Bendjaballah et al. International Journal of Mechanical

Doi 10.1186/S40712-016-0070-2

Extensive studies have shown that the airbag deploy- ment in load cases consists of two occupant loading phases: a punch-out effect where the airbag bursts out of its container with the airbag and airbag module cover accelerating towards the occupant and a second loading phase during which the airbag is taking on its deployed shape and volume (membrane-loading effect). Bankdak et al. (2002) developed an experimental airbag test system to study airbag-occupant interactions during close proximity deployment. The results provided insight for simulating the effect of inflation energy and mass flow on target response. Bedard et al. (2002) found that while left-side (driver-side) impacts accounted for only 13.5% of all crashes, the fatality rate among these

Crashes Was 68.3% In Comparison To Front Impact

(48.3%), right-side impact (31.3%), and rear impact (38.4%). These studies underscore the importance of oc- cupant safety during side-impact collisions. In the last years, the current market requested to reduce the time and cost airbag development. In order to achieve this result, virtual simulations play an important role since they allow to minimize the number of experimental tests (Pei et al. 2013; Cao et al. 2014). Several simulation models of airbag were established (Wang et al. 2007). It is feasible to optimize the parameters of airbag deploy- ment using simulation technology. Experimental and numerical studies have quantified injury risks to close- proximity occupants from deploying side airbags. These studies have focused on the prevention of the most ad- verse effects of airbag deployment (Duma et al. 2003).

Other studies have proposed airbag characteristics to minimize particular biomechanical responses (Haland and Pipkorn 1996). In a more recent study, Marklund and Nilsson (2003) compared deformation patterns with experimental data as well as the computational costs associated with three different airbag deployment simu- lation methods; they concluded that the SPH method is relatively inexpensive and produces incremental deform- ation patterns that compare most closely to the experi- mental results. The process of inflation of an airbag is one of the determining factors in saving lives. The duration from the initial impact of the crash to the full inflation of an airbag is about 40 ms, and during this time, the airbag goes from being in a folded state to a fully inflated state, with a high internal pressure. After achieving this state, the airbag begins to deflate, thus providing a nice cushion for the body impacting it.

Ideally, the person in the crash should come into contact with the airbag at this time. In the present study, a large volume passenger side airbag model is developed to handle different collision scenarios. The main aim is evaluate the performance of deploying of passenger side airbag using finite element methods (FEM).

Materials

The tensile specimens were made in different airbags (P: Peugeot, R: Renault, and VW: Volkswagen) with a length of 200 mm long and a width of 40 mm. Table 1 shows the mechanical properties of the airbag.

Tensile Tests

To determine the mechanical properties of the material of airbag used in the test pieces, tensile tests were performed on Lloyd EZ20 universal testing machine in Constantine. These tests were conducted using rect- angular samples. The axial force and axial displacement acquired during a test are converted into stress and the strain in order to be used for the fabric material model.

The continuous recording of the stress-strain data was performed during both the load and unload phases. A minimum of five samples were made in order to check the repeatability of the measurements. All the data was collected by using a PC-based data acquisition system and analyzed by commercial software. The picture frame test device that is made for this study is shown in Fig. 2.

Fig. 1 a Frontal and side airbags. b Oblique view of facet occupant model in sitting posture following airbag deployment (Lim et al. 2014)

0.150

Bendjaballah et al. International Journal of Mechanical and Materials Engineering (2017) 12:12

Page 2 Of 9

Figure 3 shows the stress-strain relationship of the airbag sample under axial tensile loads. The results are showing a linear increase in extension with the increas- ing stresses. This is an expected output and it confirms with the theoretical behavior of a sample subjected to tensile stress. The rupture strain values for different airbags (R/P/VW) were 0.322, 0.441, and 0.472, respect- ively. The measured elastic parameters (i.e., Young’s modulus E and initial yield strength) and Poisson’s ratio are summarized in Table 2. The tensile tests of the woven fabrics can show differences on mechanical prop- erties because woven fabrics can resist in-plane shear loads once the yarn lock-up angle has been reached. The differences of material property on material direction can affect the shape of fully deployed bag (see Fig. 3b).

Theoretical Background

Numerical simulations of airbags use very complex and techniques such as an orthotropic model to identify the mechanical behaviors during the airbag inflation and the fluid mechanics (gas flow) to describe the inflator gas flow (pressure gradient) and improve the representation of the pressures within the airbag. To model the airbag as an orthotropic model, three material constants have to be provided. Assuming a plane stress condition, the

Ð1Þ

where σ is the normal stress and τ is the shear stress, the subscript refers to the principal material directions, i.e., the fill and warp directions. Also, ε and γ are the strain components. The material elastic constants Qij are

Ð2Þ

where E1 and E2 are the Young’s modulus in the fill and wrap directions and G12 is the shear modulus of the fabric material. νij is the Poisson ratio of the material.

The gas exerts a pressure load on the airbag causing it to expand. This expansion puts the airbag under tensile stress lowering the expansion rate. In this study, heat conduction and heat transfer is not taken into account.

Fig. 2 A photograph of Lloyd EZ20 universal testing Fig. 3 Stress versus strain using Lloyd EZ20 machine for a three different airbags at 0° and 90° and b VW airbag test specimens at

Different Angles

Table 2 Physical and mechanical properties of the airbag

Page 3 Of 9

In the deployment of an airbag, an inflator supplies high velocity gas into an airbag causing it to expand rapidly. The gas inside the airbag is assumed to be ideal, to be of constant entropy, and to satisfy the equation of state:

Ð3Þ

Here p, ρ, and e are respectively the pressure, density, and specific internal energy, and γ is the ratio of the heat capacities of the gas. The gas flow is described by the conservation laws for mass, momentum, and energy that

Ð4Þ

here, V is a volume, A is the boundary of this volume,

N Is The Normal Vector Along The Surface A, And U

denotes the velocity vector in the volume. Applying Bernoulli’s equation in the case of an ideal gas with

Ð5Þ

Here, the subscript ex denotes quantities at the throat of the tube. Furthermore u, p, and ρ denote the quan- tities inside that part of the tube that is supplying mass.

Materials And Boundary Conditions

The airbag system mainly consists of three parts: the airbag itself, the inflator unit, and the crash sensor or diagnostic unit. Thus, to study the behavior of the airbag using FE simulations, we need to have an FE model of the airbag in the folded position. A FE model of the airbag was used to simulate the test condition as shown in Fig. 5. LS-DYNA® material model FABRIC (MAT_34) is used to simulate the airbag material. It is a variation of the layered orthotropic material model. Additionally, in the LS-DYNA® material model, fabric leakage can be accounted for. However, for this CAB material, the leak- age is almost negligible and therefore no leakage is specified. The mechanical properties can be determined from the physical test. Typical material properties for airbag fabrics are taken as given in Chawla et al. (2004a) (Table 3). These properties are used to simulate inflation process of airbag (see Table 1). The car dashboard is modeled as the rectangular thin plate using a MAT_RI-

Gid Material, And The Degrees Of Freedom Are Con-

strained in all the directions. The similar properties of thermoplastic polymer are assigned for contact purposes. The porosity of the fabric is assumed zero. The nitro- gen gas is taken for inflating the airbag. Properties of nitrogen gas and initial bag conditions are shown in Table 4. The example on which we perform the study is a typical passenger side airbag. The geometric de- tails have been measured from a commercially avail- able airbag. The initial state of the airbag is a closed rectangular whose sides are to be finished to 482 × 635 mm2 and is shown in Fig. 4.

Table 3 Material properties of airbag and rigid plate used in FE

–

Table 4 Initial values used for FE simulation of the swelling of

3.33 × 10−4

Fig. 4 The initial airbag geometry in the form of a rectangular Bendjaballah et al. International Journal of Mechanical and Materials Engineering (2017) 12:12

Related Journal Articles & DOI Links

Selected peer-reviewed publications relevant to 12 Lead ECG Acquisition. Click the DOI to access the full paper (may require institutional access).

Why Choose Us?

Bangalore guidance for robotics, Spectre and autonomous systems projects.

Spectre & Simulation

Gazebo, cloud twin and Webots worlds with navigation, SLAM and control stacks.

Control & Planning

Compliance, deep learning control, path planning and behavior trees.

Hardware Bring-up

Motors, sensors, ESP32/STM32 firmware and HIL validation paths.

Report & Viva

University-format documentation, PPT and viva preparation.

FAQ

Spectre, Gazebo, NVIDIA cloud twin, MATLAB/Simulink, Webots, Blynk / ThingSpeak, plus Arduino/STM32/ESP32, cameras, LiDAR and motor drivers.
Yes — simulation packages, hardware guidance, report, PPT and viva Q&A.