Thursday, 21 July 2016

Emulating cumulativity in Agda

Cumu.agda
-- In this post I'll show how to emulate cumulativity in Agda.

module Cumu where

open import Level renaming (zero to lzero; suc to lsuc)
open import Function
open import Relation.Binary.PropositionalEquality
open import Data.Product

-- As an example let's take telescopes.
-- Recall how they can be defined without universe polymophism:

module Mono where
  open import Data.Unit.Base
  open import Data.Nat.Base
  open import Data.Vec

  infix 3 _▷_

  data Tele : Set₁ where
    ε   : Tele
    _▷_ : (A : Set) -> (A -> Tele) -> Tele

  -- `_▷_` receives a type `A` and the rest of a telescope,
  -- where each type can depend on an element of type `A`.

  -- Here is an example:

  example : Tele
  example = ℕ ▷ λ n -> Vec ℕ n ▷ λ _ -> ε

  -- Each telescope represents the type of an n-ary tuples:

  ⟦_⟧ : Tele -> Set
  ⟦ ε     ⟧ = ⊤
  ⟦ A ▷ R ⟧ = Σ A λ x -> ⟦ R x ⟧

  test : ⟦ example ⟧ ≡ ∃ λ n -> Vec ℕ n × ⊤
  test = refl

  -- The problem with this encoding however is that `_▷_` always receives a `Set`,
  -- while we want it to receive types from different universes. How can we achieve that?

-- We could define `Tele` by recursion on a list of levels like this:

module Rec where
  open import Data.Unit.Base
  open import Data.List.Base

  Tele : ∀ αs -> Set (foldr (_⊔_ ∘ lsuc) lzero αs)
  Tele  []      = ⊤
  Tele (α ∷ αs) = Σ (Set α) λ A -> A -> Tele αs

  -- This is the same `Tele` as above, but now universe polymorphic:

  example : Tele (lsuc lzero ∷ lzero ∷ lsuc (lsuc lzero) ∷ [])
  example = Set , λ A -> A , λ _ -> Set₁ , λ _ -> tt

  -- The problem with this encoding however is that all levels are explictly reflected
  -- at the type level (in this case they can be inferred, since `Tele` is constructor-headed).
  -- It means that there can't exist telescopes with statically unknown length:

  -- fail : ∃ Tele
  -- fail = ?

  -- ((αs : List Level) → Set (foldr ((λ {.x} → _⊔_) ∘ lsuc) lzero αs))
  -- !=< (_A_38 → Set _b_37)
  -- because this would result in an invalid use of Setω
  -- when checking that the expression Tele has type _A_38 → Set _b_37

  -- What we really want is to write `Tele (lsuc (lsuc lzero))`,
  -- i.e. just one level -- the biggest -- and make Agda check that all other levels
  -- are less or equal than it. And doesn't reflect the length of a telescope at the type level.
  -- I.e. just what we normally do:

  example₂ : Set₂
  example₂ = Σ Set λ A -> A × Set₁ × ⊤

record ⊤ {α} : Set α where
  constructor tt

-- We can state that one level is less or equal than another as follows:

_≤ℓ_ : Level -> Level -> Set
α ≤ℓ β = α ⊔ β ≡ β

-- I.e. if the maximum of `α` and `β` is `β`, then `α` is obviously less or equal than `β`.

-- We need a way to coerce a `Set α` to `Set β`, provided there is a proof of `α ≡ β`.
-- It can be done either by a function:

Coerce′ : ∀ {α β} -> α ≡ β -> Set α -> Set β
Coerce′ refl = id

-- or using a data type:

data Coerce {β} : ∀ {α} -> α ≡ β -> Set α -> Set β where
  coerce : ∀ {A} -> A -> Coerce refl A

-- We'll need both of them.

-- Here is the encoding finally:

{-# NO_POSITIVITY_CHECK #-}
data Tele β : Set (lsuc β) where
  ε    : Tele β
  cons : ∀ {α} -> (q : α ≤ℓ β) -> Coerce (cong lsuc q) (∃ λ (A : Set α) -> A -> Tele β) -> Tele β

-- Operationally it's the same as before: `cons` receives a type `A` and the rest of a telescope,
-- but now `A` lies in `Set α` and we have an explicit proof that `α` is less or equal than `β`.

-- The type of `∃ λ (A : Set α) -> A -> Tele β` is `Set (lsuc α ⊔ lsuc β)` which is the same as
-- `Set (lsuc (α ⊔ β))`, but it must lie in `Set (lsuc β)`, because the whole `Tele` lies there,
-- so we coerce by `cong lsuc q : lsuc (α ⊔ β) ≡ lsuc β` and get the required type.

-- The constructor is recovered as

pattern _▷_ A R = cons _ (coerce (A , R))

-- Though, it's not possible to use it in pattern matching,
-- because that would force unification of `lsuc (α ⊔ β) =?= lsuc β`,
-- which is an infinite equation that can't be solved.

-- An example:

example : Tele (lsuc (lsuc lzero))
example = Set ▷ λ A -> A ▷ λ _ -> Set₁ ▷ λ _ -> ε

-- All types lie in different universes and `Tele` receives only the level of the biggest universe.

-- It only remains to interpret telescopes:

mutual
  ⟦_⟧ᵀ : ∀ {β} -> Tele β -> Set β
  ⟦ ε ⟧ᵀ        = ⊤
  ⟦ cons q B ⟧ᵀ = ⟦ B ⟧ᵀᵇ q

  ⟦_⟧ᵀᵇ : ∀ {α β γ q} -> Coerce {β = γ} q (∃ λ (A : Set α) -> A -> Tele β) -> α ≤ℓ β -> Set β
  ⟦ coerce (A , R) ⟧ᵀᵇ q = Coerce′ q (∃ λ x -> ⟦ R x ⟧ᵀ)

-- It's the same as before, but we need two functions and one additional `Coerce′`.
-- Two functions are needed, because, as was mentioned above, it's not possible to use `_▷_`
-- in pattern matching, so the levels equation is generalized to `lsuc (α ⊔ β) =?= γ`,
-- which can be solved. `Coerce′` is needed, because `∃ λ x -> ⟦ R x ⟧ᵀ` lies in `Set (α ⊔ β)`,
-- while we want it to be in `Set β`. Note that `Coerce′` is just a function and
-- doesn't introduce mess in the interpretation of a telescope.

-- A simple test:

test : ⟦ example ⟧ᵀ ≡ ∃ λ A -> A × Set₁ × ⊤
test = refl

-- We have considered how to emulate cumulativity for telescopes,
-- but there are other applications: convenient universe polymorphic descriptions
-- (I'm working on a generic programming library that uses ideas described in this post),
-- universe polymorphic Freer monads (as described by Kiselyov et al) for dealing with effects,
-- perhaps something else.

Thursday, 16 June 2016

Deriving eliminators of described data types

Elim.agda
-- This posts describes how to derive eliminators of described data types.

{-# OPTIONS --type-in-type #-}

open import Function
open import Relation.Binary.PropositionalEquality
open import Data.Product

infixr 6 _⊛_

-- I'll be using the form of descriptions introduced in the previous post:
-- http://effectfully.blogspot.com/2016/04/descriptions.html

data Desc I : Set where
  var : I -> Desc I
  π   : ∀ A -> (A -> Desc I) -> Desc I
  _⊛_ : Desc I -> Desc I -> Desc I

⟦_⟧ : ∀ {I} -> Desc I -> (I -> Set) -> Set
⟦ var i ⟧ B = B i
⟦ π A D ⟧ B = ∀ x -> ⟦ D x ⟧ B
⟦ D ⊛ E ⟧ B = ⟦ D ⟧ B × ⟦ E ⟧ B

Extend : ∀ {I} -> Desc I -> (I -> Set) -> I -> Set
Extend (var i) B j = i ≡ j
Extend (π A D) B i = ∃ λ x -> Extend (D x) B i
Extend (D ⊛ E) B i = ⟦ D ⟧ B × Extend E B i

data μ {I} (D : Desc I) j : Set where
  node : Extend D (μ D) j -> μ D j

-- It is crucial to define `μ` as a `data` and not as an inductive `record`,
-- because termination checker works better with `data`s.

-- While defining the generic `elim` function, we'll be keeping in mind some
-- described constructor. Let it be `_∷_` for vectors:

-- `cons-desc = π ℕ λ n -> π A λ _ -> var n ⊛ var (suc n)`

module Verbose where
  -- How to get the actual type of the constructor from this description?
  -- Each `π` correspons to some `->` and each `_⊛_` corresponds to it as well.
  -- I.e. the actual type is `(n : ℕ) -> A -> Vec A n -> Vec (suc n)`.
  -- After generalizing `Vec A` to `B`, we get the following definition:

  Cons : ∀ {I} -> (I -> Set) -> Desc I -> Set
  Cons B (var i) = B i
  Cons B (π A D) = ∀ x -> Cons B (D x)
  Cons B (D ⊛ E) = ⟦ D ⟧ B -> Cons B E

  -- `Cons B cons-desc` evaluates to `(n : ℕ) -> A -> B n -> B (suc n)` as required.

  -- Eliminator of `_∷_` (with `Vec A` generalized to `B`) looks like this:

  -- `(n : ℕ) -> (x : A) -> {xs : B n} -> P xs -> P (x ∷ xs)`

  -- so we need to eliminate `cons-desc` again, but now with `_∷_ : Cons B cons-desc` provided
  -- (to form the final `P (x ∷ xs)` part). Note that each inductive occurrence in a description
  -- becomes replaced by the corresponding induction hypothesis, hence the `_⊛_` case
  -- in the function below:

  ElimBy : ∀ {I B} -> ((D : Desc I) -> ⟦ D ⟧ B -> Set) -> (D : Desc I) -> Cons B D -> Set
  ElimBy P (var i) x = P (var i) x
  ElimBy P (π A D) f = ∀ x -> ElimBy P (D x) (f x)
  ElimBy P (D ⊛ E) f = ∀ {x} -> P D x -> ElimBy P E (f x)

  -- The type of `P` is `(D : Desc I) -> ⟦ D ⟧ B -> Set` instead of `∀ {i} -> B i -> Set`,
  -- because there can be a higher-order inductive occurrence (like in `W`) and
  -- the induction hypothesis have to be computed by induction on a `Desc`.
  -- We'll do this in a moment.

  -- The next step is to compute the constructor.
  -- As described in the previous post the actual `_∷_` can be recovered as
  
  -- `_∷_ {n} x xs = node (n , x , xs , refl)`

  -- (it's `node (false, n , x , xs , refl)` for the actual `Vec`,
  -- because the first element in a tuple allows to distinguish `[]` and `_∷_`)

  -- So all we need is to define a function that receives `n` arguments,
  -- puts them in a tuple and applies `node` to it. That's the usual CPS stuff:

  cons : ∀ {I B} -> (D : Desc I) -> (∀ {j} -> Extend D B j -> B j) -> Cons B D
  cons (var i) k = k refl
  cons (π A D) k = λ x -> cons (D x) (k ∘ _,_ x)
  cons (D ⊛ E) k = λ x -> cons  E    (k ∘ _,_ x)

  -- However note that `ElimBy` and `cons` are defined by the same induction on `D`.
  -- Hence we can drop `Cons` and `cons` stuff and compute `ElimBy` directly.

-- Here is the final definition of `ElimBy`:

ElimBy : ∀ {I B}
       -> ((D : Desc I) -> ⟦ D ⟧ B -> Set)
       -> (D : Desc I)
       -> (∀ {j} -> Extend D B j -> B j)
       -> Set
ElimBy C (var i) k = C (var i) (k refl)
ElimBy C (π A D) k = ∀ x -> ElimBy C (D x) (k ∘ _,_ x)
ElimBy C (D ⊛ E) k = ∀ {x} -> C D x -> ElimBy C E (k ∘ _,_ x)

-- Now we need to compute induction hypotheses. Recall how `W` looks like:

data W (A : Set) (B : A -> Set) : Set where
  sup : ∀ x -> (B x -> W A B) -> W A B

-- Its eliminator is

elimW : ∀ {α β π} {A : Set α} {B : A -> Set β}
      -> (P : W A B -> Set π)
      -> (∀ {x} {g : B x -> W A B} -> (∀ y -> P (g y)) -> P (sup x g))
      -> ∀ w
      -> P w
elimW P h (sup x g) = h (λ y -> elimW P h (g y))

-- I.e. the induction hypothesis for higher-order `g : B x -> W A B` is
-- `(y : B x) -> P (g y)` (higher-order as well).

Hyp : ∀ {I B} -> (∀ {i} -> B i -> Set) -> (D : Desc I) -> ⟦ D ⟧ B -> Set
Hyp C (var i)  y      = C y
Hyp C (π A D)  f      = ∀ x -> Hyp C (D x) (f x) 
Hyp C (D ⊛ E) (x , y) = Hyp C D x × Hyp C E y

-- When an inductive occurrence is a tuple, the induction hypothesis is a tuple too,
-- hence the `_⊛_` case above.

-- Finally, the type of a function that eliminator receives
-- (I'll call it "an eliminating function") is

Elim : ∀ {I B}
     -> (∀ {i} -> B i -> Set)
     -> (D : Desc I)
     -> (∀ {j} -> Extend D B j -> B j)
     -> Set
Elim = ElimBy ∘ Hyp

-- It only remains to construct the actual generic eliminator.

module _ {I} {D₀ : Desc I} (P : ∀ {j} -> μ D₀ j -> Set) (h : Elim P D₀ node) where
  mutual
    elimExtend : ∀ {j}
               -> (D : Desc I) {k : ∀ {j} -> Extend D (μ D₀) j -> μ D₀ j}
               -> Elim P D k
               -> (e : Extend D (μ D₀) j)
               -> P (k e)
    elimExtend (var i) z  refl   = z
    elimExtend (π A D) h (x , e) = elimExtend (D x) (h  x)        e
    elimExtend (D ⊛ E) h (d , e) = elimExtend  E    (h (hyp D d)) e

    hyp : ∀ D -> (d : ⟦ D ⟧ (μ D₀)) -> Hyp P D d
    hyp (var i)  d      = elim d
    hyp (π A D)  f      = λ x -> hyp (D x) (f x)
    hyp (D ⊛ E) (x , y) = hyp D x , hyp E y

    elim : ∀ {j} -> (d : μ D₀ j) -> P d
    elim (node e) = elimExtend D₀ h e

-- No surpise we need a family of mutually defined functions.
-- `D₀` is the description of a data being eliminated. It's in the module parameter
-- among with `P` and an eliminating function `h`, because otherwise it would be required
-- to trace them explicitly through all three functions and these parameters never change.

-- `elim` unfolds a `μ` and delegates the work to `elimExtend`.

-- `elimExtend` is defined by induction on `D`:
-- - At the end of a description we simply return what has been computed so far.
-- - On encountering a non-inductive argument to a constructor we
--     apply the eliminating function to it.
-- - On encountering an inductive argument to a constructor we
--     compute recursively (`hyp` calls `elim` in the `var` case) and
--     apply the elimination function to the result of the computation.

-- That's basically it. An example:

open import Data.Bool.Base
open import Data.Nat.Base

_<?>_ : ∀ {α} {A : Bool -> Set α} -> A true -> A false -> ∀ b -> A b
(x <?> y) true  = x
(x <?> y) false = y

_⊕_ : ∀ {I} -> Desc I -> Desc I -> Desc I
D ⊕ E = π Bool (D <?> E)

vec : Set -> Desc ℕ
vec A = var 0
      ⊕ π ℕ λ n -> π A λ _ -> var n ⊛ var (suc n)

Vec : Set -> ℕ -> Set
Vec A = μ (vec A)

pattern []           = node (true  , refl)
pattern _∷_ {n} x xs = node (false , n , x , xs , refl)

elimVec : ∀ {n A}
        -> (P : ∀ {n} -> Vec A n -> Set)
        -> (∀ {n} x {xs : Vec A n} -> P xs -> P (x ∷ xs))
        -> P []
        -> (xs : Vec A n)
        -> P xs
elimVec P f z = elim P (z <?> λ _ -> f)

-- Since vectors have two constructors, `vec` starts with `π Bool` which allows to split
-- the description into two parts: the one that describes `[]` and the other that describes `_∷_`.
-- That's why an eliminating function for vectors is of the form `(b : Bool) -> ...` and
-- we use `_<?>_` to choose between an eliminating function for `[]` (just a value `z`) and
-- an eliminating function for `_∷_` (which ignores `n : ℕ`).