Skip to content
MASA-Safe-RL
Parameterized PPO V2
Search
sacktock/MASA-Safe-RL
MASA-Safe-RL
sacktock/MASA-Safe-RL
Home
Get Started
Get Started
Quick Start
Core Concepts
Core Concepts
Overview
Labelling Function
Cost Function
Basic Usage
Common API
Common API
Constraints
Constraints
Overview
Constrained Markov Decision Process (CMDP)
Linear Temporal Logic (LTL) Safety Constraint
Probabilistic Computation Tree Logic (PCTL) Constraint
Step-wise Probabilistic Constraint
Reach-avoid Constraint
Multi-Agent Constraints
Multi-Agent Constraints
Overview
Constrained Markov Game (CMG)
Alternating-Time Temporal Logic (ATL) Safety
Wrappers
Wrappers
Overview
Core Wrappers
Misc Wrappers
Vectorized Envs
Metrics
Metrics
Overview
Logging
Linear Temporal Logic (LTL)
Linear Temporal Logic (LTL)
Overview
Propositional Formula
DFA
Cost Function as a DFA
Shaped Cost Function
Probabilistic Computation Tree Logic (PCTL)
Probabilistic Computation Tree Logic (PCTL)
Overview
Propositional Operators
Temporal Operators
Model Checking
Environments
Environments
Multi Agent
Multi Agent
Overview
Gridworlds
Gridworlds
Overview
Capture the Flag
Clean Up
Markov Stag Hunt
Matrix Games
Matrix Games
Overview
Bertrand
Chicken
Congestion
Dynamic Public Goods Game
Inspection
Single Agent
Single Agent
Overview
Cartpole
Safety Gridworlds
Safety Gridworlds
Overview
Island Navigation
Conveyor Belt
Sokoban
Pacman
Gridworlds
Gridworlds
Overview
Bridge Crossing
Colour Grid World
Colour Bomb Grid World
Media Streaming
Algorithms
Algorithms
Algorithms Overview
Tabular Algorithms
Tabular Algorithms
Overview
Q Learning
Q Learning Lambda
LCRL
SEM
RECREG
Multi-agent Tabular Algorithms
Multi-agent Tabular Algorithms
Overview
IQL
IQL Lagrangian
IQL Lambda
On-Policy Algorithms
On-Policy Algorithms
Overview
A2C
PPO
PPO Lagrangian
Constrained Policy Optimization
Shielding
Shielding
Probabilistic Shielding
Winning-region safety-game shielding
Shielded Algorithms
Shielded Algorithms
Overview
Multi-agent coalition safety-game shielding
Parameterized PPO
Parameterized PPO V2
Tutorials
Tutorials
Basics
Basics
Overview
First MASA Experiment
Labels, Costs, and Infos
Wrapper Stack
Constraints
Constraints
Overview
Constraints Tour
LTL-Safety
LTL-Safety
Overview
LTL Safety Colour Bomb
Wrappers
Wrappers
Overview
Vectorization and Normalization
Environments
Environments
Overview
Create a New Environment
Baselines
Baselines
Overview
Tabular Safe RL Baselines
Continuous Safe RL Baselines
Shielding
Shielding
Overview
Probabilistic Shielding MiniPacman
Safety Abstractions Pacman Coins
FrozenLake Shielding
Multi-Agent
Multi-Agent
Overview
Multi-Agent CMG
Parameterized PPO V2
¶
Source:
masa/prob_shield/parameterized_ppo_v2.py
Back to top