Random vector functional link networks for function approximation on manifolds

Needell, Deanna; Nelson, Aaron A.; Saab, Rayan; Salanevich, Palina; Schavemaker, Olov

doi:10.3389/fams.2024.1284706

ORIGINAL RESEARCH article

Front. Appl. Math. Stat., 17 April 2024

Sec. Optimization

Volume 10 - 2024 | https://doi.org/10.3389/fams.2024.1284706

This article is part of the Research TopicOptimization for Low-rank Data Analysis: Theory, Algorithms and ApplicationsView all articles

Random vector functional link networks for function approximation on manifolds

Deanna Needell¹^*

¹Department of Mathematics, University of California, Los Angeles, Los Angeles, CA, United States
²Department of Mathematical Sciences, United States Air Force Academy, Colorado Springs, CO, United States
³Department of Mathematics and Halıcıoğlu Data Science Institute, University of California, San Diego, San Diego, CA, United States
⁴Mathematical Institute, Utrecht University, Utrecht, Netherlands

The learning speed of feed-forward neural networks is notoriously slow and has presented a bottleneck in deep learning applications for several decades. For instance, gradient-based learning algorithms, which are used extensively to train neural networks, tend to work slowly when all of the network parameters must be iteratively tuned. To counter this, both researchers and practitioners have tried introducing randomness to reduce the learning requirement. Based on the original construction of Igelnik and Pao, single layer neural-networks with random input-to-hidden layer weights and biases have seen success in practice, but the necessary theoretical justification is lacking. In this study, we begin to fill this theoretical gap. We then extend this result to the non-asymptotic setting using a concentration inequality for Monte-Carlo integral approximations. We provide a (corrected) rigorous proof that the Igelnik and Pao construction is a universal approximator for continuous functions on compact domains, with approximation error squared decaying asymptotically like O(1/n) for the number n of network nodes. We then extend this result to the non-asymptotic setting, proving that one can achieve any desired approximation error with high probability provided n is sufficiently large. We further adapt this randomized neural network architecture to approximate functions on smooth, compact submanifolds of Euclidean space, providing theoretical guarantees in both the asymptotic and non-asymptotic forms. Finally, we illustrate our results on manifolds with numerical experiments.

1 Introduction

In recent years, neural networks have once again triggered an increased interest among researchers in the machine learning community. So-called deep neural networks model functions using a composition of multiple hidden layers, each transforming (possibly non-linearly) the previous layer before building a final output representation [1–5]. In machine learning parlance, these layers are determined by sets of weights and biases that can be tuned so that the network mimics the action of a complex function. In particular, a single layer feed-forward neural network (SLFN) with n nodes may be regarded as a parametric function $f_{n} : ℝ^{N} \to ℝ$ of the form

f_{n} (x) = \sum_{k = 1}^{n} v_{k} ρ (〈 w_{k}, x 〉 + b_{k}), x \in ℝ^{N} .

Here, the function ρ:ℝ → ℝ is called an activation function and is potentially non-linear. Some typical examples include the sigmoid function $ρ (z) = \frac{1}{1 + exp (- z)}$ , ReLU ρ(z) = max{0, z}, and sign functions, among many others. The parameters of the SLFN are the number of nodes n ∈ ℕ in the the hidden layer, the input-to-hidden layer weights and biases ${w_{k}}_{k = 1}^{n} \subset ℝ^{N}$ and ${b_{k}}_{k = 1}^{n} \subset ℝ$ (resp.), and the hidden-to-output layer weights ${v_{k}}_{k = 1}^{n} \subset ℝ$ . In this way, neural networks are fundamentally parametric families of functions whose parameters may be chosen to approximate a given function.

It has been shown that any compactly supported continuous function can be approximated with any given precision by a single layer neural network with a suitably chosen number of nodes [6], and harmonic analysis techniques have been used to study stability of such approximations [7]. Other recent results that take a different approach directly analyze the capacity of neural networks from a combinatorial point of view [8, 9].

While these results ensure existence of a neural network approximating a function, practical applications require construction of such an approximation. The parameters of the neural network can be chosen using optimization techniques to minimize the difference between the network and the function f:ℝ^N → ℝ it is intended to model. In practice, the function f is usually not known, and we only have access to a set ${(x_{k}, f (x_{k}))}_{k = 1}^{m}$ of values of the function at finitely many points sampled from its domain, called a training set. The approximation error can be measured by comparing the training data to the corresponding network outputs when evaluated on the same set of points, and the parameters of the neural network f_n can be learned by minimizing a given loss function $L (x_{1}, \dots, x_{m})$ ; a typical loss function is the sum-of-squares error

L (x_{1}, \dots, x_{m}) = \frac{1}{m} \sum_{k = 1}^{m} | f (x_{k}) - f_{n} (x_{k}) |^{2} .

The SLFN which approximates f is then determined using an optimization algorithm, such as back-propagation, to find the network parameters which minimize $L (x_{1}, \dots, x_{m})$ . It is known that there exist weights and biases which make the loss function vanish when the number of nodes n is at least m, provided the activation function is bounded, non-linear, and has at least one finite limit at either ±∞ [10].

Unfortunately, optimizing the parameters in SLFNs can be difficult. For instance, any non-linearity in the activation function can cause back-propagation to be very time-consuming or get caught in local minima of the loss function [11]. Moreover, deep neural networks can require massive amounts of training data, and so are typically unreliable for applications with very limited data availability, such as agriculture, healthcare, and ecology [12].

To address some of the difficulties associated with training deep neural networks, both researchers and practitioners have attempted to incorporate randomness in some way. Indeed, randomization-based neural networks that yield closed form solutions typically require less time to train and avoid some of the pitfalls of traditional neural networks trained using back-propagation [11, 13, 14]. One of the popular randomization-based neural network architectures is the Random Vector Functional Link (RVFL) network [15, 16], which is a single layer feed-forward neural network in which the input-to-hidden layer weights and biases are selected randomly and independently from a suitable domain and the remaining hidden-to-output layer weights are learned using training data.

By eliminating the need to optimize the input-to-hidden layer weights and biases, RVFL networks turn supervised learning into a purely linear problem. To see this, define ρ(X) ∈ ℝ^n×m to be the matrix whose jth column is ${ρ (〈 w_{k}, x_{j} 〉 + b_{k})}_{k = 1}^{n}$ and f(X) ∈ ℝ^m the vector whose jth entry is f(x_j). Then, the vector v ∈ ℝⁿ of hidden-to-output layer weights is the solution to the matrix-vector equation f(X) = ρ(X)^Tv, which can be solved by computing the Moore-Penrose pseudoinverse of ρ(X)^T. In fact, there exist weights and biases that make the loss function vanish when the number of nodes n is at least m, provided the activation function is smooth [17].

Although originally considered in the early- to mid-1990s [15, 16, 18, 19], RVFL networks have had much more recent success in several modern applications, including time-series data prediction [20], handwritten word recognition [21], visual tracking [22], signal classification [23, 24], regression [25], and forecasting [26, 27]. Deep neural network architectures based on RVFL networks have also made their way into more recent literature [28, 29], although traditional, single layer RVFL networks tend to perform just as well as, and with lower training costs than, their multi-layer counterparts [29].

Even though RVFL networks are proving their usefulness in practice, the supporting theoretical framework is currently lacking [see 30]. Most theoretical research into the approximation capabilities of deep neural networks centers around two main concepts: universal approximation of functions on compact domains and point-wise approximation on finite training sets [17]. For instance, in the early 1990s, it was shown that multi-layer feed-forward neural networks having activation functions that are continuous, bounded, and non-constant are universal approximators (in the L^p sense for 1 ≤ p < ∞) of continuous functions on compact domains [31, 32]. The most notable result in the existing literature regarding the universal approximation capability of RVFL networks is due to Igelnik and Pao [16] in the mid-1990s, who showed that such neural networks can universally approximate continuous functions on compact sets; the noticeable lack of results since has left a sizable gap between theory and practice. In this study, we begin to bridge this gap by further improving the Igelnik and Pao result, and bringing the mathematical theory behind RFVL networks into the modern spotlight. Below, we introduce the notation that will be used throughout this study, and describe our main contributions.

1.1 Notation

For a function f:ℝ^N → ℝ, the set supp(f)⊂ℝ^N denotes the support of f. We denote by $C_{c} (ℝ^{N})$ and $C_{0} (ℝ^{N})$ the classes of continuous functions mapping ℝ^N to ℝ whose support sets are compact and vanish at infinity, respectively. Given a set S⊂ℝ^N, we define its radius to be $rad (S) : = {sup}_{x \in S} ∥ x ∥_{2}$ ; moreover, if dμ denotes the uniform volume measure on S, then we write $vol (S) : = \int_{S} d μ$ to represent the volume of S. For any probability distribution P:ℝ^N → [0, 1], a random variable X distributed according to P is denoted by X~P, and we write its expectation as $𝔼 X : = \int_{ℝ^{N}} X d P$ . The open ℓ_p ball of radius r > 0 centered at x ∈ ℝ^N is denoted by $B_{p}^{N} (x, r)$ for all 1 ≤ p ≤ ∞; the ℓ_p unit-ball centered at the origin is abbreviated $B_{p}^{N}$ . Given a fixed δ > 0 and a set S⊂ℝ^N, a minimal δ-net for S, which we denote $C (δ, S)$ , is the smallest subset of S satisfying $S \subset \cup_{x \in C (δ, S)} B_{2}^{N} (x, δ)$ ; the δ-covering number of S is the cardinality of a minimal δ-net for S and is denoted $N (δ, S) : = | C (δ, S) |$ .

1.2 Main results

In this study, we analyze the uniform approximation capabilities of RVFL networks. More specifically, we consider the problem of using RVFL networks to estimate a continuous, compactly supported function on N-dimensional Euclidean space.

The first theoretical result on approximating properties of RVFL networks, due to Igelnik and Pao [16], guarantees that continuous functions may be universally approximated on compact sets using RVFL networks, provided the number of nodes n ∈ ℕ in the network goes to infinity. Moreover, it shows that the mean square error of the approximation vanishes at a rate proportional to 1/n. At the time, this result was state-of-the-art and justified how RVFL networks were used in practice. However, the original theorem is not technically correct. In fact, several aspects of the proof technique are flawed. Some of the minor flaws are mentioned in Li et al. [33], but the subsequent revisions do not address the more significant issues which would make the statement of the result technically correct. We address these issues in this study, see Remark 1. Thus, our first contribution to the theory of RVFL networks is a corrected version of the original Igelnik and Pao theorem:

Theorem 1 ([16]). Let $f \in C_{c} (ℝ^{N})$ with K: = supp(f) and fix any activation function ρ, such that either ρ ∈ L¹(ℝ)∩L^∞(ℝ) with $\int_{ℝ} ρ (z) d z = 1$ or ρ is differentiable with ρ′ ∈ L¹(ℝ)∩L^∞(ℝ) and $\int_{ℝ} ρ^{'} (z) d z = 1$ . For any ε > 0, there exist distributions from which input weights ${w_{k}}_{k = 1}^{n}$ and biases ${b_{k}}_{k = 1}^{n}$ are drawn, and there exist hidden-to-output layer weights ${v_{k}}_{k = 1}^{n} \subset ℝ$ that depend on the realization of weights and biases, such that the sequence of RVFL networks ${f_{n}}_{n = 1}^{\infty}$ is defined by

f_{n} (x) : = \sum_{k = 1}^{n} v_{k} ρ (〈 w_{k}, x 〉 + b_{k}) for x \in K

satisfies

𝔼 \int_{K} | f (x) - f_{n} (x) |^{2} d x < ε + O (1 / n),

as n → ∞.

For a more precise formulation of Theorem 1 and its proof, we refer the reader to Theorem 5 and Section 3.1.

Remark 1.

1. Even though in Theorem 1 we only claim existence of the distribution for input weights ${w_{k}}_{k = 1}^{n}$ and biases ${b_{k}}_{k = 1}^{n}$ , such a distribution is actually constructed in the proof. Namely, for any ε > 0, there exist constants α, Ω > 0 such that the random variables

\begin{array}{l} w_{k} ~ Unif ({[- α Ω, α Ω]}^{N}); \\ y_{k} ~ Unif (K); \\ u_{k} ~ Unif ([- \frac{π}{2} (2 L + 1), \frac{π}{2} (2 L + 1)]), \\ where L : = ⌈ \frac{2N}{π} rad (K) Ω - \frac{1}{2} ⌉, \end{array}

are independently drawn from their associated distributions, and b_k: = −〈w_k, y_k〉 −αu_k.

2. We note that, unlike the original theorem statement in Igelnik and Pao [16], Theorem 1 does not show exact convergence of the sequence of constructed RVFL networks f_n to the original function f. Indeed, it only ensures that the limit f_n is ε-close to f. This should still be sufficient for practical applications since, given a desired accuracy level ε > 0, one can find values of α, Ω, andn such that this accuracy level is achieved on average. Exact convergence can be proved if one replaces α and Ω in the distribution described above by sequences ${α_{n}}_{n = 1}^{\infty}$ and ${Ω_{n}}_{n = 1}^{\infty}$ of positive numbers, both tending to infinity with n. In this setting, however, there is no guaranteed rate of convergence; moreover, as n increases, the ranges of the random variables ${w_{k}}_{k = 1}^{n}$ and ${u_{k}}_{k = 1}^{n}$ become increasingly larger, which may cause problems in practical applications.

3. The approach we take to construct the RVFL network approximating a function f allows one to compute the output weights ${v_{k}}_{k = 1}^{n}$ exactly (once the realization of random parameters is fixed), in the case where the function f is known. For the details, we refer the reader to Equations 6, 8 in the proof of Theorem 1. If we only have access to a training set that is sufficiently large and uniformly distributed over the support of f, these formulas can be used to compute the output weights approximately, instead of solving the least squares problem.

4. Note that the normalization $\int_{ℝ} ρ (z) d z = 1$ of the activation function can be replaced by the condition $\int_{ℝ} ρ (z) d z \neq 0$ . Indeed, in the case when ρ ∈ L¹(ℝ)∩L^∞(ℝ) and $\int_{ℝ} ρ (z) d z \notin {0, 1},$ one can simply use Theorem 1 to approximate $\frac{1}{\int_{ℝ} ρ (z) d z} f$ by a sequence of RVFL network with the activation function $\frac{1}{\int_{ℝ} ρ (z) d z} ρ$ . Mutatis mutandis in the case when $\int_{ℝ} ρ^{'} (z) d z^{'} \notin {0, 1} .$ More generally, this trick allows any of our theorems to be applied in the case $\int_{ℝ} ρ (z) d z \neq 0 .$

One of the drawbacks of Theorem 1 is that the mean square error guarantee is asymptotic in the number of nodes used in the neural network. This is clearly impractical for applications, and so it is desirable to have a more explicit error bound for each fixed number n of nodes used. To this end, we provide a new, non-asymptotic version of Theorem 1, which provides an error guarantee with high probability whenever the number of network nodes is large enough, albeit at the price of an additional Lipschitz requirement on the activation function:

Theorem 2 Let $f \in C_{c} (ℝ^{N})$ with K: = supp(f) and fix any activation function ρ ∈ L¹(ℝ)∩L^∞(ℝ) with $\int_{ℝ} ρ (z) d z = 1 .$ Suppose further that ρ is κ-Lipschitz on ℝ for some κ > 0. For any ε > 0 and η ∈ (0, 1), suppose that n ≥ C(N, f, ρ)ε⁻¹log(η⁻¹/ε), where C(N, f, ρ) is independent of ε and η and depends on f, ρ, and superexponentially on N. Then, there exist distributions from which input weights ${w_{k}}_{k = 1}^{n}$ and biases ${b_{k}}_{k = 1}^{n}$ are drawn, and there exist hidden-to-output layer weights ${v_{k}}_{k = 1}^{n} \subset ℝ$ that depend on the realization of weights and biases, such that the RVFL network defined by

f_{n} (x) : = \sum_{k = 1}^{n} v_{k} ρ (〈 w_{k}, x 〉 + b_{k}) for x \in K

satisfies

\int_{K} | f (x) - f_{n} (x) |^{2} d x < ε

with probability at least 1−η.

For simplicity, the bound on the number n of the nodes on the hidden layer here is rough. For a more precise formulation of this result that contains a bound with explicit constant, we refer the reader to Theorem 6 in Section 3.2. We also note that the distribution of the input weight and bias here can be selected as described in Remark 1.

The constructions of RVFL networks presented in Theorems 1, 2 depend heavily on the dimension of the ambient space ℝ^N. If N is small, this dependence does not present much of a problem. However, many modern applications require the ambient dimension to be large. Fortunately, a common assumption in practice is the support of the signals of interest lies on a lower-dimensional manifold embedded in ℝ^N. For instance, the landscape of cancer cell states can be modeled using non-linear, locally continuous “cellular manifolds;” indeed, while the ambient dimension of this state space is typically high (e.g., single-cell RNA sequencing must account for approximately 20,000 gene dimensions), cellular data actually occupies an intrinsically lower dimensional space [34]. Similarly, while the pattern space of neural population activity in the brain is described by an exponential number of parameters, the spatiotemporal dynamics of brain activity lie on a lower-dimensional subspace or “neural manifold” [35]. In this study, we propose a new RVFL network architecture for approximating continuous functions defined on smooth compact manifolds that allows to replace the dependence on the ambient dimension N with dependence on the manifold intrinsic dimension. We show that RVFL approximation results can be extended to this setting. More precisely, we prove the following analog of Theorem 2.

Theorem 3 Let $M \subset ℝ^{N}$ be a smooth, compact d-dimensional manifold with finite atlas {(_{U_j, ϕ_j)}j ∈ J} and $f \in C (M)$ . Fix any activation function ρ ∈ L¹(ℝ)∩L^∞(ℝ) with $\int_{ℝ} ρ (z) d z = 1$ such that ρ is κ-Lipschitz on ℝ for some κ > 0. For any ε > 0 and η ∈ (0, 1), suppose n ≥ C(d, f, ρ)ε⁻¹log(η⁻¹/ε), where C(d, f, ρ) is independent of ε and η and depends on f, ρ, and superexponentially on d. Then, there exists an RVFL-like approximation f_n of the function f with a parameter selection similar to the Theorem 1 construction that satisfies

\int_{M} | f (x) - f_{n} (x) |^{2} d x < ε

with probability at least 1−η.

For a the construction of the RVFL-like approximation f_n and a more precise formulation of this result and an analog of Theorem 1 applied to manifolds, we refer the reader to Section 3.3.1 and Theorems 7, 8. We note that the approximation f_n here is not obtained as a single RVFL network construction, but rather as a combination of several RVFL networks in local manifold coordinates.

1.3 Organization

The remaining part of the article is organized as follows. In Section 2, we discuss some theoretical preliminaries on concentration bounds for Monte-Carlo integration and on smooth compact manifolds. Monte-Carlo integration is an essential ingredient in our construction of RVFL networks approximating a given function, and we use the results listed in this section to establish approximation error bounds. Theorem 1 is proven in Section 3.1, where we break down the proof into four main steps, constructing a limit-integral representation of the function to be approximated in Lemmas 3, 4, then using Monte-Carlo approximation of the obtained integral to construct an RVFL network in Lemma 5, and, finally, establishing approximation guarantees for the constructed RVFL network. The proofs of Lemmas 3, 4, and 5 can be found in Sections 3.5.1, 3.5.2, and 3.5.3, respectively. We further study properties of the constructed RVFL networks and prove the non-asymptotic approximation result of Theorem 2 in Section 3.2. In Section 3.3, we generalize our results and propose a new RVFL network architecture for approximating continuous functions defined on smooth compact manifolds. We show that RVFL approximation results can be extended to this setting by proving an analog of Theorem 1 in Section 3.3.2 and Theorem 3 in Section 3.5.5. Finally, in Section 3.4, we provide numerical evidence to illustrate the result of Theorem 3.

2 Materials and methods

In this section, we briefly introduce supporting material and theoretical results which we will need in later sections. This material is far from exhaustive, and is meant to be a survey of definitions, concepts, and key results.

2.1 A concentration bound for classic Monte-Carlo integration

A crucial piece of the proof technique employed in Igelnik and Pao [16], which we will use repeatedly, is the use of the Monte-Carlo method to approximate high-dimensional integrals. As such, we start with the background on Monte-Carlo integration. The following introduction is adapted from the material in Dick et al. [36].

Let f:ℝ^N → ℝ and S⊂ℝ^N a compact set. Suppose we want to estimate the integral $I (f, S) : = \int_{S} f d μ$ , where μ is the uniform measure on S. The classic Monte Carlo method does this by an equal-weight cubature rule,

I_{n} (f, S) : = \frac{vol (S)}{n} \sum_{j = 1}^{n} f (x_{j}),

where ${x_{j}}_{j = 1}^{n}$ are independent identically distributed uniform random samples from S and $vol (S) : = \int_{S} d μ$ is the volume of S. In particular, note that 𝔼I_n(f, S) = I(f, S) and

𝔼 I_{n} {(f, S)}^{2} = \frac{1}{n} vol (S) I (f^{2}, S) + (n - 1) I {(f, S)}^{2} .

Let us define the quantity

\begin{array}{l} σ {(f, S)}^{2} : = \frac{I (f^{2}, S)}{vol (S)} - \frac{I {(f, S)}^{2}}{vo l^{2} (S)} . & (1) \end{array}

It follows that the random variable I_n(f) has mean I(f, S) and variance vol²(S)σ(f, S)²/n. Hence, by the Central Limit Theorem, provided that 0 < vol²(S)σ(f, S)² < ∞, we have

\lim_{n \to \infty} ℙ (| I_{n} (f, S) - I (f, S) | \leq \frac{C ε (f, S)}{\sqrt{n}}) = {(2 π)}^{- 1 / 2} \int_{- C}^{C} e^{- x^{2} / 2} d x

for any constant C > 0, where ε(f, S): = vol(S)σ(f, S). This yields the following well-known result:

Theorem 4 For any f ∈ L²(S, μ), the mean-square error of the Monte Carlo approximation I_n(f, S) satisfies

𝔼 {| I}_{n} (f, S) - I (f, S) |^{2} = \frac{vo l^{2} (S) σ {(f, S)}^{2}}{n},

where the expectation is taken with respect to the random variables ${x_{j}}_{j = 1}^{n}$ and σ(f, S) is defined in Equation 1.

In particular, Theorem 4 implies $𝔼 I_{n} (f, S) - I (f, S)^{2} = O (1 / n)$ as n → ∞.

In the non-asymptotic setting, we are interested in obtaining a useful bound on the probability ℙ(|I_n(f, S)−I(f, S)| ≥ t) for all t > 0. The following lemma follows from a generalization of Bennett's inequality (Theorem 7.6 in [37]; see also [38, 39]).

Lemma 1 For any f ∈ L²(S) and n ∈ ℕ, we have

ℙ (| I_{n} (f, S) - I (f, S) | \geq t) \leq 3 exp (- \frac{n t}{C K} log (1 + \frac{K t}{vol (S) I (f^{2}, S)}))

for all t > 0 and a universal constant C > 0, provided |vol(S)f(x)| ≤ K for almost every x ∈ S.

2.2 Smooth, compact manifolds in Euclidean space

In this section, we review several concepts of smooth manifolds that will be useful to us later. Many of the definitions and results that follow can be found, for instance, in Shaham et al. [40]. Let $M \subset ℝ^{N}$ be a smooth, compact d-dimensional manifold. A chart for $M$ is a pair (U, ϕ) such that $U \subset M$ is an open set and ϕ:U → ℝ^d is a homeomorphism. One way to interpret a chart is as a tangent space at some point x ∈ U; in this way, a chart defines a Euclidean coordinate system on U via the map ϕ. A collection {(_{U_j, ϕ_j)}j ∈ J} of charts defines an atlas for $M$ if $\cup_{j \in J} U_{j} = M$ . We now define a special collection of functions on $M$ called a partition of unity.

Definition 1 Let $M \subset ℝ^{N}$ be a smooth manifold. A partition of unity of $M$ with respect to an open cover {_{U_j}j ∈ J} of $M$ is a family of non-negative smooth functions {_{η_j}j ∈ J} such that for every $x \in M$ , we have $1 = \sum_{j \in J} η_{j} (x)$ and, for every j ∈ J, supp(η_j)⊂U_j.

It is known that if $M$ is compact, there exists a partition of unity of $M$ such that supp(η_j) is compact for all j ∈ J [see 41]. In particular, such a partition of unity exists for any open cover of $M$ corresponding to an atlas.

Fix an atlas {(_{U_j, ϕ_j)}j ∈ J} for $M$ , as well as the corresponding, compactly supported partition of unity {_{η_j}j ∈ J}. Then, we have the following useful result [see 40, Lemma 4.8].

Lemma 2 Let $M \subset ℝ^{N}$ be a smooth, compact manifold with atlas {(_{U_j, ϕ_j)}j ∈ J} and compactly supported partition of unity {_{η_j}j ∈ J}. For any $f \in C (M)$ , we have

f (x) = \sum_{{j \in J : x \in U_{j}}} ({\hat{f}}_{j} ◦ ϕ_{j}) (x)

for all $x \in M$ , where

{\hat{f}}_{j} (z) : = {\begin{array}{l} f (ϕ_{j}^{- 1} (z)) η_{j} (ϕ_{j}^{- 1} (z)) & z \in ϕ_{j} (U_{j}) \\ 0 & otherwise . \end{array}

In later sections, we use the representation of Lemma 2 to integrate functions $f \in C (M)$ over $M$ . To this end, for each j ∈ J, let Dϕ_j(y) denote the differential of ϕ_j at y ∈ U_j, which is a map from the tangent space $T_{y} M$ into ℝ^d. One may interpret Dϕ_j(y) as the matrix representation of a basis for the cotangent space at y ∈ U_j. As a result, Dϕ_j(y) is necessarily invertible for each y ∈ U_j, and so we know that |det(Dϕ_j(y))| > 0 for each y ∈ U_j. Hence, it follows by the change of variables theorem that

\begin{array}{l} \int_{ℳ} f (x) d x = \int_{ℳ} \sum_{{j \in J : x \in U_{j}}} ({\hat{f}}_{j} ° ϕ_{j}) (x) d x \\ = \sum_{j \in J} \int_{ϕ_{j} (U_{j})} \frac{{\hat{f}}_{j} (z)}{| \det (D ϕ_{j} (ϕ_{j}^{- 1} (z))) |} d z . & (2) \end{array}

3 Results

In this section, we prove our main results formulated in Section 1.2 and also use numerical simulations to illustrate the RVFL approximation performance in a low-dimensional submanifold setup. To improve readability of this section, we postpone the proofs of technical lemmas till Section 3.5.

3.1 Proof of Theorem 1

We split the proof of the theorem into two parts, first handling the case ρ ∈ L¹(ℝ)∩L^∞(ℝ) and second, addressing the case ρ′ ∈ L¹(ℝ)∩L^∞(ℝ).

3.1.1 Proof of Theorem 1 when ρ ∈ L¹(ℝ)∩L^∞(ℝ)

We begin by restating the theorem in a form that explicitly includes the distributions that we draw our random variables from.

Theorem 5 ([16]) Let $f \in C_{c} (ℝ^{N})$ with K: = supp(f) and fix any activation function ρ ∈ L¹(ℝ)∩L^∞(ℝ) with $\int_{ℝ} ρ (z) d z = 1$ . For any ε > 0, there exist constants α, Ω > 0 such that the following holds: If, for k ∈ ℕ, the random variables

\begin{array}{l} w_{k} ~ Unif ({[- α Ω, α Ω]}^{N}); \\ y_{k} ~ Unif (K); \\ u_{k} ~ Unif ([- \frac{π}{2} (2 L + 1), \frac{π}{2} (2 L + 1)]), \\ where L : = ⌈ \frac{2N}{π} rad (K) Ω - \frac{1}{2} ⌉, \end{array}

are independently drawn from their associated distributions, and

b_{k} : = - 〈 w_{k}, y_{k} 〉 - α u_{k},

then there exist hidden-to-output layer weights ${v_{k}}_{k = 1}^{n} \subset ℝ$ (that depend on the realization of the weights ${w_{k}}_{k = 1}^{n}$ and biases ${b_{k}}_{k = 1}^{n}$ ) such that the sequence of RVFL networks ${f_{n}}_{n = 1}^{\infty}$ defined by

f_{n} (x) : = \sum_{k = 1}^{n} v_{k} ρ (〈 w_{k}, x 〉 + b_{k}) for x \in K

satisfies

𝔼 \int_{K} | f (x) - f_{n} (x) |^{2} d x \leq ε + O (1 / n) .

as n → ∞.

Proof. Our proof technique is based on that introduced by Igelnik and Pao and can be divided into four steps. The first three steps essentially consist of Lemma 3, Lemma 4, and Lemma 5, and the final step combines them to obtain the desired result. First, the function f is approximated by a convolution, given in Lemma 3. The proof of this result can be found in Section 3.5.1.

Lemma 3 Let $f \in C_{0} (ℝ^{N})$ and h ∈ L¹(ℝ^N) with $\int_{ℝ^{N}} h (z) d z = 1 .$ For Ω > 0, define

\begin{array}{l} h_{Ω} (y) : = Ω^{N} h (Ω y) . & (3) \end{array}

Then, we have

\begin{array}{l} f (x) = lim_{Ω \to \infty} (f * h_{Ω}) (x) & (4) \end{array}

uniformly for all x ∈ ℝ^N.

Next, we represent f as the limiting value of a multidimensional integral over the parameter space. In particular, we replace (f*h_Ω)(x) in the convolution identity (Equation 4) with a function of the form $\int_{K} F (y) ρ (〈 w, x 〉 + b (y)) d y$ , as this will introduce the RVFL structure we require. To achieve this, we first use a truncated cosine function in place of the activation function ρ and then switch back to a general activation function.

To that end, for each fixed Ω > 0, let $L = L (Ω) : = ⌈ \frac{2 N}{π} rad (K) Ω - \frac{1}{2} ⌉$ and define cos_Ω:ℝ → [−1, 1] by

\begin{array}{l} \cos_{Ω} (x) : = {\begin{array}{l} \cos (x) & x \in [- \frac{1}{2} (2 L + 1) π, \frac{1}{2} (2 L + 1) π], \\ 0 & otherwise . \end{array} & (5) \end{array}

Moreover, introduce the functions

\begin{array}{l} F_{α, Ω} (y, w, u) : = \frac{α}{{(2 π)}^{N}} f (y) \cos_{Ω} (u) \prod_{j = 1}^{N} ϕ (w (j) / Ω), \\ b_{α} (y, w, u) : = - α (〈 w, y 〉 + u) & (6) \end{array}

where y, w ∈ ℝ^N, u ∈ ℝ, and ϕ = A*A for any even function A ∈ C^∞(ℝ) supported on $[- \frac{1}{2}, \frac{1}{2}]$ s.t. ∥A∥₂ = 1. Then, we have the following lemma, a detailed proof of which can be found in Section 3.5.2.

Lemma 4 Let $f \in C_{c} (ℝ^{N})$ and ρ ∈ L¹(ℝ) with K: = supp(f) and $\int_{ℝ} ρ (z) d z = 1$ . Define F_{α, Ω} and b_α as in Equation 6 for all α > 0. Then, for $L : = ⌈ \frac{2 N}{π} rad (K) Ω - \frac{1}{2} ⌉$ , we have

\begin{array}{l} f (x) = lim_{Ω \to \infty} lim_{α \to \infty} \int_{K (Ω)} F_{α, Ω} (y, w, u) ρ (α 〈 w, x 〉 + b_{α} (y, w, u)) d y d w d u & (7) \end{array}

uniformly for every x ∈ K, where $K (Ω) : = K \times {[- Ω, Ω]}^{N} \times [- \frac{π}{2} (2 L + 1), \frac{π}{2} (2 L + 1)]$ .

The next step in the proof of Theorem 5 is to approximate the integral in Equation 7 using the Monte-Carlo method. Define $v_{k} : = \frac{vol (K (Ω))}{n} F_{α, Ω} y_{k}, \frac{w_{k}}{α}, u_{k}$ for k = 1, …, n, and the random variables ${f_{n}}_{n = 1}^{\infty}$ by

\begin{array}{l} f_{n} (x) : = \sum_{k = 1}^{n} v_{k} ρ (〈 w_{k}, x 〉 + b_{k}) . & (8) \end{array}

Then, we have the following lemma that is proven in Section 3.5.3.

Lemma 5 Let $f \in C_{c} (ℝ^{N})$ and ρ ∈ L¹(ℝ)∩L^∞(ℝ) with K: = supp(f) and $\int_{ℝ} ρ (z) d z = 1$ . Then, as n → ∞, we have

\begin{array}{l} 𝔼 \int_{K} | \int_{K (Ω)} F_{α, Ω} (y, w, u) ρ (α 〈 w, x 〉 + b_{α} (y, w, u)) d y d w d u \\ {- f_{n} (x) |}^{2} d x = O (1 / n), & (9) \end{array}

where $K (Ω) : = K \times {[- Ω, Ω]}^{N} \times [- \frac{π}{2} (2 L + 1), \frac{π}{2} (2 L + 1)]$ and $L : = ⌈ \frac{2 N}{π} rad (K) Ω - \frac{1}{2} ⌉$ .

To complete the proof of Theorem 5, we combine the limit representation (Equation 7) with the Monte-Carlo error guarantee (Equation 9) and show that, given any ε > 0, there exist α, Ω > 0 such that

𝔼 \int_{K} | f (x) - f_{n} (x) |^{2} d x \leq ε + O (1 / n)

as n → ∞. To this end, let ε′ > 0 be arbitrary and consider the integral I(x; p) given by

\begin{array}{l} I (x; p) : = \int_{K (Ω)} {(F}_{α, Ω} (y, w, u) ρ (α 〈 w, x 〉 + b_{α} (y, w, u)))^{p} d y d w d u & (10) \end{array}

for x ∈ K and p ∈ ℕ. By Equation 7, there exist α, Ω > 0 such that |f(x)−I(x; 1)| < ε′ holds for every x ∈ K, and so it follows that

| f (x) - f_{n} (x) | < ε^{'} + | I (x; 1) - f_{n} (x) |

for every x ∈ K. Jensen's inequality now yields that

\begin{array}{l} 𝔼 \int_{K} | f (x) - f_{n} (x) |^{2} d x \leq 2 vol (K) {(ε^{'})}^{2} + 2 𝔼 \int_{K} I (x; 1) - f_{n} (x)^{2} d x . & (11) \end{array}

By Equation 9, we know that the second term on the right-hand side of Equation 11 is O(1/n). Therefore, we have

𝔼 \int_{K} | f (x) - f_{n} (x) |^{2} d x \leq 2 vol (K) {(ε^{'})}^{2} + O (1 / n),

and so the proof is completed by taking $ε^{'} = \sqrt{ε / 2 vol (K)}$ and choosing α, Ω > 0 accordingly.

3.1.2 Proof of Theorem 1 when ρ′ ∈ L¹(ℝ)∩L^∞(ℝ)

The full statement of the theorem is identical to that of Theorem 5 albeit now with ρ′ ∈ L¹(ℝ)∩L^∞(ℝ), so we omit it for brevity. Its proof is also similar to the proof of the case where ρ ∈ L¹(ℝ)∩L^∞(ℝ) with some key modifications. Namely, one uses an integration by parts argument to modify the part of the proof corresponding to Lemma 4. The details of this argument are presented in Section 3.5.4.

3.2 Proof of Theorem 2

In this section, we prove the non-asymptotic result for RVFL networks in ℝ^N, and we begin with a more precise statement of the theorem that makes all the dimensional dependencies explicit.

Theorem 6 Consider the hypotheses of Theorem 5 and suppose further that ρ is κ-Lipschitz on ℝ for some κ > 0. For any

0 < δ < \frac{\sqrt{ε}}{8 \sqrt{2 N} κ α^{2} M Ω {(Ω / π)}^{N} vo l^{3 / 2} (K) (π + 2 N rad (K) Ω)},

suppose

n \geq \frac{c Σ α {(Ω / π)}^{N} (π + 2 N rad (K) Ω) log (3 η^{- 1} N (δ, K))}{\sqrt{ε} log (1 + \frac{\sqrt{ε}}{Σ α {(Ω / π)}^{N} (π + 2 N rad (K) Ω)})},

where $M : = sup_{x \in K} | f (x) |$ , c > 0 is a numerical constant, and Σ is a constant depending on f and ρ, and let parameters ${w_{k}}_{k = 1}^{n}$ , ${b_{k}}_{k = 1}^{n}$ , and ${v_{k}}_{k = 1}^{n}$ be as in Theorem 5. Then, the RVFL network defined by

f_{n} (x) : = \sum_{k = 1}^{n} v_{k} ρ (〈 w_{k}, x 〉 + b_{k}) for x \in K

satisfies

\int_{K} | f (x) - f_{n} (x) |^{2} d x < ε

with probability at least 1−η.

Proof. Let $f \in C_{c} (ℝ^{N})$ with K: = supp(f) and suppose ε > 0, η ∈ (0, 1) are fixed. Take an arbitrarily κ-Lipschitz activation function ρ ∈ L¹(ℝ)∩L^∞(ℝ). We wish to show that there exists an RVFL network ${f_{n}}_{n = 1}^{\infty}$ defined on K that satisfies the

\int_{K} | f (x) - f_{n} (x) |^{2} d x < ε

with probability at least 1−η when n is chosen sufficiently large. The proof is obtained by modifying the proof of Theorem 5 for the asymptotic case.

We begin by repeating the first two steps in the proof of Theorem 5 from Sections 3.5.1, 3.5.2. In particular, by Lemma 4 we have the representation given by Equation 4, namely,

f (x) = lim_{Ω \to \infty} lim_{α \to \infty} \int_{K (Ω)} F_{α, Ω} (y, w, u) ρ (α 〈 w, x 〉 + b_{α} (y, w, u)) d y d w d u

holds uniformly for all x ∈ K. Hence, if we define the random variables f_n and I_n from Section 3.5.3 as in Equations 8, 29, respectively, we seek a uniform bound on the quantity

| f (x) - f_{n} (x) | \leq | f (x) - I (x; 1) | + | I_{n} (x) - I (x; 1) |

over the compact set K, where I(x; 1) is given by Equation 10 for all x ∈ K. Since Equation 7 allows us to fix α, Ω > 0 such that

\begin{array}{l} | f (x) - I (x; 1) | = | f (x) - \\ \int_{K (Ω)} F_{α, Ω} (y, w, u) ρ (α 〈 w, x 〉 + b_{α} (y, w, u)) d y d w d u | < \sqrt{\frac{ε}{2 vol (K)}} \end{array}

holds for every x ∈ K simultaneously, the result would follow if we show that, with high probability, $| I_{n} (x) - I (x; 1) | < \sqrt{ε / 2 vol (K)}$ uniformly for all x ∈ K. Indeed, this would yield

\begin{array}{l} \int_{K} | f (x) - f_{n} (x) |^{2} d x \leq 2 \int_{K} | f (x) - I (x; 1) |^{2} d x \\ + 2 \int_{K} | I_{n} (x) - I (x; 1) |^{2} d x < ε \end{array}

with high probability. To this end, for δ > 0, let $C (δ, K) \subset K$ denote a minimal δ-net for K, with cardinality $N (δ, K)$ . Now, fix x ∈ K and consider the inequality

\begin{array}{l} | I_{n} (x) - I (x; 1) | \leq \underset{(*)}{\underset{︸}{| I_{n} (x) - I_{n} (z) |}} + \underset{(* *)}{\underset{︸}{| I_{n} (z) - I (z; 1) |}} \\ + \underset{(***)}{\underset{︸}{| I (x; 1) - I (z; 1) |}}, & (12) \end{array}

where $z \in C (δ, K)$ is such that ∥x−z∥₂ < δ. We will obtain the desired bound on Equation 12 by bounding each of the terms (*), (**), and (***) separately.

First, we consider the term (*). Recalling the definition of I_n, observe that we have

\begin{array}{l} (*) = \frac{vol (K (Ω))}{n} | \sum_{k = 1}^{n} F_{α, Ω} (y_{k}, w_{k}, u_{k}) (ρ (α 〈 w_{k}, x 〉 + b_{α} (y_{k}, w_{k}, u_{k})) \\ - ρ (α 〈 w_{k}, z 〉 + b_{α} (y_{k}, w_{k}, u_{k}))) | \\ \leq \frac{α M vol (K (Ω))}{{(2 π)}^{N} n} \sum_{k = 1}^{n} | ρ (α 〈 w_{k}, x 〉 + b_{α} (y_{k}, w_{k}, u_{k})) \\ - ρ (α 〈 w_{k}, z 〉 + b_{α} (y_{k}, w_{k}, u_{k})) | \\ \leq α M {(2 π)}^{- N} vol (K (Ω)) R_{α, Ω} (x, z), \end{array}

where $M : = sup_{x \in K} | f (x) |$ and we define

\begin{array}{l} R_{α, Ω} (x, z) : = \sup_{\begin{matrix} y \in K \\ w \in {[- Ω, Ω]}^{N} \\ u \in [- (L + \frac{1}{2}) π, (L + \frac{1}{2}) π] \end{matrix}} | ρ (α 〈 w, x 〉 + b_{α} (y, w, u)) \\ - ρ (α 〈 w, z 〉 + b_{α} (y, w, u)) | . \end{array}

Now, since ρ is assumed to be κ-Lipschitz, we have

\begin{array}{l} | ρ (α 〈 w, x 〉 + b_{α} (y, w, u)) - ρ (α 〈 w, z 〉 + b_{α} (y, w, u)) | \\ = | ρ (α (〈 w, x - y 〉 - u)) \\ - ρ (α (〈 w, z - y 〉 - u)) | \leq κ α | 〈 w, x - z 〉 | \end{array}

for any y ∈ K, w ∈ [−Ω, Ω]^N, and $u \in [- (L + \frac{1}{2}) π, (L + \frac{1}{2}) π] .$ Hence, an application of the Cauchy–Schwarz inequality yields $R_{α, Ω} (x, z) \leq κ α Ω δ \sqrt{N}$ for all x ∈ K, from which it follows that

\begin{array}{l} (*) \leq M \sqrt{N} κ δ α^{2} Ω {(2 π)}^{- N} vol (K (Ω)) & (13) \end{array}

holds for all x ∈ K.

Next, we bound (***) using a similar approach. Indeed, by the definition of I(·;1), we have

\begin{array}{l} (***) = | \int_{K (Ω)} F_{α, Ω} (y, w, u) (ρ (α 〈 w, x 〉 + b_{α} (y, w, u)) \\ - ρ (α 〈 w, z 〉 + b_{α} (y, w, u))) d y d w d u | \\ \leq \frac{α M ‖ ϕ ‖_{\infty}^{N}}{{(2 π)}^{N}} \int_{K (Ω)} | ρ (α 〈 w, x 〉 + b_{α} (y, w, u)) \\ - ρ (α 〈 w, z 〉 + b_{α} (y, w, u)) | d y d w d u \\ \leq α M {(2 π)}^{- N} vol (K (Ω)) R_{α, Ω} (x, z) . \end{array}

Using the fact that $R_{α, Ω} (x, z) \leq κ α Ω δ \sqrt{N}$ for al x ∈ K, it follows that

\begin{array}{l} (***) \leq M \sqrt{N} κ δ α^{2} Ω {(2 π)}^{- N} vol (K (Ω)) & (14) \end{array}

holds for all x ∈ K, just like Equation 13.

Notice that the Equations 13, 14 are deterministic. In fact, both can be controlled by choosing an appropriate value for δ in the net $C (δ, K)$ . To see this, fix ε′ > 0 arbitrarily and recall that vol(K(Ω)) = (2Ω)^Nπ(2L + 1)vol(K). A simple computation then shows that (*) + (***) < ε′ whenever

\begin{array}{l} δ & < \frac{ε^{'}}{4 \sqrt{N} κ α^{2} M Ω {(Ω / π)}^{N} vol (K) (π + 2 N rad (K) Ω)} & (15) \end{array}

\begin{array}{l} < \frac{ε^{'}}{2 \sqrt{N} κ α^{2} M Ω {(Ω / π)}^{N} π (2 L + 1) vol (K)} . \end{array}

We now bound (**) uniformly for x ∈ K. Unlike (*) and (***), we cannot bound this term deterministically. In this case, however, we may apply Lemma 1 to

g_{z} (y, w, u) : = F_{α, Ω} (y, w, u) ρ (α 〈 w, z 〉 + b_{α} (y, w, u)),

for any $z \in C (δ, K)$ . Indeed, $g_{z} \in L^{2} (K (Ω))$ because $F_{α, Ω} \in L^{2} (K (Ω))$ and ρ ∈ L^∞(ℝ). Then, Lemma 1 yields the tail bound

\begin{array}{l} ℙ ((* *) \geq t) = ℙ (| I_{n} (g_{z}, K (Ω)) - I (g_{z}, K (Ω)) | t \geq t) \\ \leq 3 \exp (- \frac{n t}{B c} \log (1 + \frac{B t}{vol (K (Ω)) I (g_{z}^{2}, K (Ω))})) \\ = 3 \exp (- \frac{n t}{B c} \log (1 + \frac{B t}{vol (K (Ω)) I (z; 2)})) \end{array}

for all t > 0, where c > 0 is a numerical constant and

\begin{array}{l} B : = 2 α M {(Ω / π)}^{N} (π + 2 N r a d (K) Ω) ∥ ρ ∥_{\infty} v o l (K) \\ \geq α M {(Ω / π)}^{N} π (2 L + 1) ∥ ρ ∥_{\infty} v o l (K) \\ = α M {(2 π)}^{- N} ∥ ρ ∥_{\infty} v o l (K (Ω)) \\ \geq \max_{z \in C (δ, K)} ∥ g_{z} ∥_{\infty} v o l (K (Ω)) . \end{array}

By taking

C : = 2 M ∥ ρ ∥_{\infty} vol (K) and Σ : = 2 C \sqrt{2 vol (K)},

we obtain B = Cα(Ω/π)^N(π+2Nrad(K)Ω) and

max_{z \in C (δ, K)} vol (K (Ω)) I (z; 2) \leq (α M {(2 π)}^{- N} ∥ ρ ∥_{\infty} vol (K (Ω)))^{2} \leq B^{2} .

If we choose the number of nodes such that

\begin{array}{l} n \geq \frac{B c log (3 η^{- 1} N (δ, K))}{t log (1 + t / B)}, & (16) \end{array}

then a union bound yields (**) < t simultaneously for all $z \in C (δ, K)$ with probability at least 1−η. Combined with the bounds from Equations 13, 14, it follows from Equation 12 that

| I_{n} (x) - I (x; 1) | < ε^{'} + t

simultaneously for all x ∈ K with probability at least 1−η, provided δ and n satisfy Equations 15, 16, respectively. Since we require $| I_{n} (x) - I (x; 1) | < \sqrt{ε / 2 vol (K)}$ , the proof is then completed by setting $ε^{'} + t = \sqrt{ε / 2 vol (K)}$ and choosing δ and n accordingly. In particular, it suffices to choose $ε^{'} = t = \frac{1}{2} \sqrt{ε / 2 vol (K)} = C \sqrt{ε} / Σ,$ so that Equations 15, 16 become

\begin{array}{l} δ & < \frac{\sqrt{ε}}{8 \sqrt{2 N} κ α^{2} M Ω {(Ω / π)}^{N} vo l^{3 / 2} (K) (π + 2 N rad (K) Ω)}, \\ n & \geq \frac{c Σ α {(Ω / π)}^{N} (π + 2 N rad (K) Ω) log (3 η^{- 1} N (δ, K))}{\sqrt{ε} log (1 + \frac{\sqrt{ε}}{Σ α {(Ω / π)}^{N} (π + 2 N rad (K) Ω)})}, \end{array}

as desired.

Remark 2 The implication of Theorem 6 is that, given a desired accuracy level ε > 0, one can construct a RVFL network f_n that is ε-close to f with high probability, provided the number of nodes n in the neural network is sufficiently large. In fact, if we assume that the ambient dimension N is fixed here, then δ and n depend on the accuracy ε and probability η as

δ ≲ \sqrt{ε} and n ≳ \frac{log (η^{- 1} N (δ, K))}{\sqrt{ε} log (1 + \sqrt{ε})} .

Using that log(1+x) = x+O(x²) for small values of x, the requirement on the number of nodes behaves like

n ≳ \frac{log (η^{- 1} N (\sqrt{ε}, K)}{ε}

whenever ε is sufficiently small. Using a simple bound on the covering number, this yields a coarse estimate of n≳ε⁻¹log(η⁻¹/ε).

Remark 3 If we instead assume that N is variable, then, under the assumption that f is Hölder continuous with exponent β, one should expect that n = ω(N^2βN) as N → ∞ (in light of Remark 10 and in conjunction with Theorem 6 with log(1+1/x)≈1/x for large x). In other words, the number of nodes required in the hidden layer is superexponential in the dimension. This dependence of n on N may be improved by means of more refined proof techniques. As for α, if follows from Remark 12 that α = Θ(1) as N → ∞ provided $\int_{ℝ} | v ρ (v) | d v < \infty .$

Remark 4 The κ-Lipschitz assumption on the activation function ρ may likely be removed. Indeed, since (***) in Equation 12 can be bounded instead by leveraging continuity of the L¹ norm with respect to translation, the only term whose bound depends on the Lipschitz property of ρ is (*). However, the randomness in I_n (that we did not use to obtain the bound in Equation 13) may be enough to control (*) in most cases. Indeed, to bound (*), we require control over quantities of the form |ρ(α(〈w_k, x−y_k〉 −u_k−ρα〈w_k, z−y_k〉−u_k))|. For most practical realizations of ρ, this difference will be small with high probability (on the draws of y_k, w_k, u_k), whenever ∥x−z∥₂ is sufficiently small.

3.3 Results on sub-manifolds of Euclidean space

The constructions of RVFL networks presented in Theorems 5, 6 depend heavily on the dimension of the ambient space ℝ^N. Indeed, the random variables used to construct the input-to-hidden layer weights and biases for these neural networks are N-dimensional objects; moreover, it follows from Equations 15, 16 that the lower bound on the number n of nodes in the hidden layer depends superexponentially on the ambient dimension N. If the ambient dimension is small, these dependencies do not present much of a problem. However, many modern applications require the ambient dimension to be large. Fortunately, a common assumption in practice is that signals of interest have (e.g., manifold) structure that effectively reduces their complexity. Good theoretical results and algorithms in a number of settings typically depend on this induced smaller dimension rather than the ambient dimension. For this reason, it is desirable to obtain approximation results for RVFL networks that leverage the underlying structure of the signal class of interest, namely, the domain of $f \in C_{c} (ℝ^{N})$ .

One way to introduce lower-dimensional structure in the context of RVFL networks is to assume that supp(f) lies on a subspace of ℝ^N. More generally, and motivated by applications, we may consider the case where supp(f) is actually a submanifold of ℝ^N. To this end, for the remainder of this section, we assume $M \subset ℝ^{N}$ to be a smooth, compact d-dimensional manifold and consider the problem of approximating functions $f \in C (M)$ using RVFL networks. As we are going to see, RVFL networks in this setting yield theoretical guarantees that replace the dependencies of Theorems 5, 6 on the ambient dimension N with dependencies on the manifold dimension d. Indeed, one should expect that the random variables ${w_{k}}_{k = 1}^{n}$ , ${b_{k}}_{k = 1}^{n}$ are essentially d-dimensional objects (rather than N-dimensional) and that the lower bound on the number of network nodes in Theorem 6 scales as a (superexponential) function of d rather than N.

3.3.1 Adapting RVFL networks to d-manifolds

As in Section 2.2, let {(_{U_j, ϕ_j)}j ∈ J} be an atlas for the smooth, compact d-dimensional manifold $M \subset ℝ^{N}$ with the corresponding compactly supported partition of unity {_{η_j}j ∈ J}. Since $M$ is compact, we assume without loss of generality that |J| < ∞. Indeed, if we additionally assume that $M$ satisfies the property that there exists an r > 0 such that, for each $x \in M$ , $M \cap B_{2}^{N} (x, r)$ is diffeomorphic to an ℓ₂ ball in ℝ^d with diffeomorphism close to the identity, then one can choose an atlas {(_{U_j, ϕ_j)}j ∈ J} with $| J | ≲ 2^{d} T_{d} vol (M) r^{- d}$ by intersecting $M$ with ℓ₂ balls in ℝ^N of radii r/2 [40]. Here, T_d is the so-called thickness of the covering and there exist coverings such that T_d≲dlog(d).

Now, for $f \in C (M)$ , Lemma 2 implies that

\begin{array}{l} f (x) = \sum_{{j \in J : x \in U_{j}}} ({\hat{f}}_{j} ◦ ϕ_{j}) (x) & (17) \end{array}

for all $x \in M$ , where

{\hat{f}}_{j} (z) : = {\begin{array}{l} f (ϕ_{j}^{- 1} (z)) η_{j} (ϕ_{j}^{- 1} (z)) & z \in ϕ_{j} (U_{j}) \\ 0 & otherwise . \end{array}

As we will see, the fact that $M$ is smooth and compact implies ${\hat{f}}_{j} \in C_{c} (ℝ^{d})$ for each j ∈ J, and so we may approximate each ${\hat{f}}_{j}$ using RVFL networks on ℝ^d as in Theorems 5, 6. In this way, it is reasonable to expect that f can be approximated on $M$ using a linear combination of these low-dimensional RVFL networks. More precisely, we propose approximating f on $M$ via the following process:

1. For each j ∈ J, approximate ${\hat{f}}_{j}$ uniformly on $ϕ_{j} (U_{j}) \subset ℝ^{d}$ using a RVFL network ${\tilde{f}}_{n_{j}}$ as in Theorems 5, 6;

2. Approximate f uniformly on $M$ by summing these RVFL networks over J, i.e.,

\begin{array}{l} f (x) \approx \sum_{{j \in J : x \in U_{j}}} ({\tilde{f}}_{n_{j}} ◦ ϕ_{j}) (x) \end{array}

for all $x \in M$ .

3.3.2 Main results on d-manifolds

We now prove approximation results for the manifold RVFL network architecture described in Section 3.3.1. For notational clarity, from here onward, we use $lim_{{n_{j}}_{j \in J} \to \infty}$ to denote the limit as each n_j tends to infinity simultaneously. The first theorem that we prove is an asymptotic approximation result for continuous functions on manifolds using the RVFL network construction presented in Section 3.3.1. This theorem is the manifold-equivalent of Theorem 5.

Theorem 7 Let $M \subset ℝ^{N}$ be a smooth, compact d-dimensional manifold with finite atlas {(_{U_j, ϕ_j)}j ∈ J} and $f \in C (M)$ . Fix any activation function ρ ∈ L¹(ℝ)∩L^∞(ℝ) with $\int_{ℝ} ρ (z) d z = 1$ . For any ε > 0, there exist constants α_j, Ω_j > 0 for each j ∈ J such that the following holds. If, for each j ∈ J and for k ∈ ℕ, the random variables

\begin{array}{l} w_{k}^{(j)} ~ Unif ({[- α_{j} Ω_{j}, α_{j} Ω_{j}]}^{d}); \\ y_{k}^{(j)} ~ Unif (ϕ_{j} (U_{j})); \\ u_{k}^{(j)} ~ Unif ([- \frac{π}{2} (2 L_{j} + 1), \frac{π}{2} (2 L_{j} + 1)]), \\ where L_{j} : = ⌈ \frac{2d}{π} rad (ϕ_{j} (U_{j})) Ω_{j} - \frac{1}{2} ⌉, \end{array}

are independently drawn from their associated distributions, and

b_{k}^{(j)} : = - 〈 w_{k}^{(j)}, y_{k}^{(j)} 〉 - α_{j} u_{k}^{(j)},

then there exist hidden-to-output layer weights ${v_{k}^{(j)}}_{k = 1}^{n_{j}} \subset ℝ$ such that the sequences of RVFL networks ${{\tilde{f}}_{n_{j}}}_{n_{j} = 1}^{\infty}$ defined by

{\tilde{f}}_{n_{j}} (z) : = \sum_{k = 1}^{n_{j}} v_{k}^{(j)} ρ (〈 w_{k}^{(j)}, z 〉 + b_{k}^{(j)}), for z \in ϕ_{j} (U_{j})

satisfy

𝔼 \int_{ℳ} | f (x) - \sum_{{j \in J : x \in U_{j}}} ({\tilde{f}}_{n_{j}} ° ϕ_{j}) (x) |^{2} d x \leq ε + O (1 / \min_{_{j \in J}} n_{j})

as {_{n_j}j ∈ J} → ∞.

Proof. We wish to show that there exist sequences of RVFL networks ${{\tilde{f}}_{n_{j}}}_{n_{j} = 1}^{\infty}$ defined on ϕ_j(U_j) for each j ∈ J, which together satisfy the asymptotic error bound

𝔼 \int_{ℳ} | f (x) - \sum_{{j \in J : x \in U_{j}}} ({\tilde{f}}_{n_{j}} ° ϕ_{j}) (x) |^{2} d x \leq ε + O (1 / \min_{_{j \in J}} n_{j})

as {_{n_j}j ∈ J} → ∞. We will do so by leveraging the result of Theorem 5 on each $ϕ_{j} (U_{j}) \subset ℝ^{d}$ .

To begin, recall that we may apply the representation given by Equation 17 for f on each chart (U_j, ϕ_j); the RVFL networks ${\tilde{f}}_{n_{j}}$ we seek are approximations of the functions ${\hat{f}}_{j}$ in this expansion. Now, as supp(η_j)⊂U_j is compact for each j ∈ J, it follows that each set ϕ_j(supp(η_j)) is a compact subset of ℝ^d. Moreover, because ${\hat{f}}_{j} (z) \neq 0$ if and only if z ∈ ϕ_j(U_j) and $ϕ_{j}^{- 1} (z) \in supp (η_{j}) \subset U_{j}$ , we have that ${\hat{f}}_{j} = {\hat{f}}_{j} |_{ϕ_{j} (supp (η_{j})}$ is supported on a compact set. Hence, ${\hat{f}}_{j} \in C_{c} (ℝ^{d})$ for each j ∈ J, and so we may apply Lemma 4 to obtain the uniform limit representation given by Equation 7 on ϕ_j(U_j), that is,

\begin{array}{l} {\hat{f}}_{j} (z) = lim_{Ω_{j} \to \infty} lim_{α_{j} \to \infty} \int_{K (Ω_{j})} F_{α_{j}, Ω_{j}} (y, w, u) ρ {(α}_{j} 〈 w, z 〉 \\ + b_{α_{j}} (y, w, u)) d y d w d u, \end{array}

where we define

K (Ω_{j}) : = ϕ_{j} (U_{j}) \times {[- Ω_{j}, Ω_{j}]}^{d} \times [- \frac{π}{2} (2 L_{j} + 1), \frac{π}{2} (2 L_{j} + 1)] .

In this way, the asymptotic error bound that is the final result of Theorem 5, namely

\begin{array}{l} 𝔼 \int_{ϕ_{j} (U_{j})} | {\hat{f}}_{j} (z) - {\tilde{f}}_{n_{j}} (z) |^{2} d z \leq ε_{j} + O (1 / n_{j}) & (18) \end{array}

holds. With these results in hand, we may now continue with the main body of the proof.

Since the representation given by Equation 17 for f on each chart (U_j, ϕ_j) yields

| f (x) - \sum_{{j \in J : x \in U_{j}}} ({\tilde{f}}_{n_{j}} ° ϕ_{j}) (x) | \leq \sum_{{j \in J : x \in U_{j}}} | ({\hat{f}}_{j} ° ϕ_{j}) (x) - ({\tilde{f}}_{n_{j}} ° ϕ_{j}) (x) |

for all $x \in M$ , Jensen's inequality allows us to bound the mean square error of our RVFL approximation by

\begin{array}{l} 𝔼 \int_{ℳ} | f (x) - \sum_{{j \in J : x \in U_{j}}} ({\tilde{f}}_{n_{j}} ° ϕ_{j}) (x) |^{2} d x \\ \leq | J | \cdot \underset{(*)}{\underset{︸}{𝔼 \int_{ℳ} \sum_{{j \in J : x \in U_{j}}} | ({\hat{f}}_{j} ° ϕ_{j}) (x) - ({\tilde{f}}_{n_{j}} ° ϕ_{j}) (x) |^{2} d x}} & (19) \end{array}

To bound (*), note that the change of variables given by Equation 2 implies

\begin{array}{l} \int_{ℳ} \sum_{{j \in J : x \in U_{j}}} | ({\hat{f}}_{j} ° ϕ_{j}) (x) - ({\tilde{f}}_{n_{j}} ° ϕ_{j}) (x) |^{2} d x \\ = \sum_{j \in J} \int_{ϕ_{j} (U_{j})} \frac{| {\hat{f}}_{j} (z) - {\tilde{f}}_{n_{j}} (z) |^{2}}{| \det (D ϕ_{j} (ϕ_{j}^{- 1} (z))) |} d z \end{array}

for each j ∈ J. Defining $β_{j} : = inf_{y \in U_{j}} | det (D ϕ_{j} (y)) |$ , which is necessarily bounded away from zero for each j ∈ J by compactness of $M$ , we therefore have

(*) \leq \sum_{j \in J} β_{j}^{- 1} 𝔼 \int_{ϕ_{j} (U_{j})} | {\hat{f}}_{j} (z) - {\tilde{f}}_{n_{j}} (z) |^{2} d z .

Hence, applying Equation 18 for each j ∈ J yields

\begin{array}{l} (*) \leq \sum_{j \in J} β_{j}^{- 1} {(ε}_{j} + O (1 / n_{j})) = \sum_{j \in J} \frac{ε_{j}}{β_{j}} + O (1 / min_{j \in J} n_{j}) & (20) \end{array}

because $\sum_{j \in J} 1 / n_{j} \leq | J | / min_{j \in J} n_{j} .$ With the bound given by Equation 20 in hand, Equation 19 becomes

𝔼 \int_{ℳ} | f (x) - \sum_{_{\begin{matrix} {j \in J : \\ x \in U_{j}} \end{matrix}}} ({\tilde{f}}_{n_{j}} ° ϕ_{j}) (x) |^{2} d x \leq | J | \sum_{j \in J} \frac{ε_{j}}{β_{j}} + O (1 / \min_{j \in J} n_{j})

as {_{n_j}j ∈ J} → ∞, and so the proof is completed by taking each ε_j > 0 in such a way that

ε = | J | \sum_{j \in J} \frac{ε_{j}}{β_{j}},

and choosing α_j, Ω_j > 0 accordingly for each j ∈ J.

Remark 5 Note that the neural-network architecture obtained in Theorem 7 has the following form in the case of a generic atlas. To obtain the estimate of f(x), the input x is first “pre-processed” by computing ϕ_j(x) for each j ∈ J such that x ∈ U_j, and then put through the corresponding RVFL network. However, using the Geometric Multi-Resolution Analysis approach from Allard et al. [42] (as we do in Section 3.4), one can construct an approximation (in an appropriate sense) of the atlas, with maps ϕ_j being linear. In this way, the pre-processing step can be replaced by the layer computing ϕ_j(x), followed by the RVFL layer f_j. We refer the reader to Section 3.4 for the details.

The biggest takeaway from Theorem 7 is that the same asymptotic mean-square error behavior we saw in the RVFL network architecture of Theorem 5 holds for our RVFL-like construction on manifolds, with the added benefit that the input-to-hidden layer weights and biases are now d-dimensional random variables rather than N-dimensional. Provided the size of the atlas |J| is not too large, this significantly reduces the number of random variables that must be generated to produce a uniform approximation of $f \in C (M)$ .

One might expect to see a similar reduction in dimension dependence for the non-asymptotic case if the RVFL network construction of Section 3.3.1 is used. Indeed, our next theorem, which is the manifold-equivalent of Theorem 6, makes this explicit:

Theorem 8 Let $M \subset ℝ^{N}$ be a smooth, compact d-dimensional manifold with finite atlas {(_{U_j, ϕ_j)}j ∈ J} and $f \in C (M)$ . Fix any activation function ρ ∈ L¹(ℝ)∩L^∞(ℝ) such that ρ is κ-Lipschitz on ℝ for some κ > 0 and $\int_{ℝ} ρ (z) d z = 1$ . For any ε > 0, there exist constants α_j, Ω_j > 0 for each j ∈ J such that the following holds. Suppose, for each j ∈ J and for k = 1, ..., n_j, the random variables

\begin{array}{l} w_{k}^{(j)} ~ Unif ({[- α_{j} Ω_{j}, α_{j} Ω_{j}]}^{d}); \\ y_{k}^{(j)} ~ Unif (ϕ_{j} (U_{j})); \\ u_{k}^{(j)} ~ Unif ([- \frac{π}{2} (2 L_{j} + 1), \frac{π}{2} (2 L_{j} + 1)]), \\ where L_{j} : = ⌈ \frac{2d}{π} rad (ϕ_{j} (U_{j})) Ω_{j} - \frac{1}{2} ⌉, \end{array}

are independently drawn from their associated distributions, and

b_{k}^{(j)} : = - 〈 w_{k}^{(j)}, y_{k}^{(j)} 〉 - α_{j} u_{k}^{(j)} .

Then, there exist hidden-to-output layer weights ${v_{k}^{(j)}}_{k = 1}^{n_{j}} \subset ℝ$ such that, for any

\begin{array}{l} 0 < δ_{j} < \\ \frac{\sqrt{ε}}{8 | J | \sqrt{d vol (M)} κ α_{j}^{2} M_{j} Ω_{j} {(Ω_{j} / π)}^{d} vol (ϕ_{j} (U_{j})) (π + 2 d rad (ϕ_{j} (U_{j})) Ω)}, \end{array}

and

\begin{array}{l} n_{j} \geq \\ \frac{2 c | J | \sqrt{vol (ℳ)} C^{(j)} α_{j} {(Ω_{j} / π)}^{d} (π + 2 d rad (ϕ_{j} (U_{j})) Ω_{j}) \log (3 | J | η^{- 1} N (δ_{j}, ϕ_{j} (U_{j})))}{\sqrt{ε} \log (1 + \frac{\sqrt{ε}}{2 | J | \sqrt{vol (ℳ)} C^{(j)} α_{j} {(Ω_{j} / π)}^{d} (π + 2 d rad (ϕ_{j} (U_{j})) Ω_{j})})}, \end{array}

where $M_{j} : = sup_{z \in ϕ_{j} (U_{j})} | {\hat{f}}_{j} (z) |$ , c > 0 is a numerical constant, and $C^{(j)} : = 2 M_{j} ∥ ρ ∥_{\infty} vol (ϕ_{j} (U_{j})),$ the sequences of RVFL networks ${{\tilde{f}}_{n_{j}}}_{n_{j} = 1}^{\infty}$ defined by

\begin{array}{l} {\tilde{f}}_{n_{j}} (z) : = \sum_{k = 1}^{n_{j}} v_{k}^{(j)} ρ (〈 w_{k}^{(j)}, z 〉 + b_{k}^{(j)}), for z \in ϕ_{j} (U_{j}) \end{array}

satisfy

\int_{ℳ} | f (x) - \sum_{{j \in J : x \in U_{j}}} ({\tilde{f}}_{n_{j}} ° ϕ_{j}) (x) |^{2} d x < ε

with probability at least 1−η.

Proof. See Section 3.5.5.

As alluded to earlier, an important implication of Theorems 7, 8 is that the random variables ${w_{k}^{(j)}}_{k = 1}^{n_{j}}$ and ${b_{k}^{(j)}}_{k = 1}^{n_{j}}$ are d-dimensional objects for each j ∈ J. Moreover, bounds for δ_j and n_j now have superexponential dependence on the manifold dimension d instead of the ambient dimension N. Thus, introducing the manifold structure removes the dependencies on the ambient dimension, replacing them instead with the intrinsic dimension of $M$ and the complexity of the atlas {(_{U_j, ϕ_j)}j ∈ J}.

Remark 6 The bounds on the covering radii δ_j and hidden layer nodes n_j needed for each chart in Theorem 8 are not optimal. Indeed, these bounds may be further improved if one uses the local structure of the manifold, through quantities such as its curvature and reach. In particular, the appearance of |J| in both bounds may be significantly improved upon if the manifold is locally well-behaved.

3.4 Numerical simulations

In this section, we provide numerical evidence to support the result of Theorem 8. Let $M \subset ℝ^{N}$ be a smooth, compact d-dimensional manifold. Since having access to an atlas for $M$ is not necessarily practical, we assume instead that we have a suitable approximation to $M$ . For our purposes, we will use a Geometric Multi-Resolution Analysis (GMRA) approximation of $M$ (see [42]; and also, e.g., [43] for a complete definition).

A GMRA approximation of $M$ provides a collection ${(C_{j}, P_{j})}_{j \in {1, \dots J}}$ of centers $C_{j} = {c_{j, k}}_{k = 1}^{K_{j}} \subset ℝ^{N}$ and affine projections $P_{j} = {P_{j, k}}_{k = 1}^{K_{j}}$ on ℝ^N such that, for each j ∈ {1, …, J}, the pairs ${(c_{j, k}, P_{j, k})}_{k = 1}^{K_{j}}$ define d-dimensional affine spaces that approximate $M$ with increasing accuracy in the following sense. For every $x \in M$ , there exists ${\tilde{C}}_{x} > 0$ and $k^{'} \in {1, \dots, K_{j}}$ such that

\begin{array}{l} ∥ x - P_{j, k^{'}} x ∥_{2} \leq {\tilde{C}}_{x} 2^{- j} & (21) \end{array}

holds whenever $∥ x - c_{j, k^{'}} ∥_{2}$ is sufficiently small.

In this way, a GMRA approximation of $M$ essentially provides a collection of approximate tangent spaces to $M$ . Hence, a GMRA approximation having fine enough resolution (i.e., large enough j) is a good substitution for an atlas. In practice, one must often first construct a GMRA from empirical data, assumed to be sampled from appropriate distributions on the manifold. Indeed, this is possible, and yields the so-called empirical GMRA, studied in Maggioni et al. [44], where finite-sample error bounds are provided. The main point is that given enough samples on the manifold, one can construct a good GMRA approximation of the manifold.

Let ${(c_{j, k}, P_{j, k})}_{k = 1}^{K_{j}}$ be a GMRA approximation of $M$ for refinement level j. Since the affine spaces defined by (c_j,k, P_j,k) for each k ∈ {1, …, K_j} are d-dimensional, we will approximate f on $M$ by projecting it (in an appropriate sense) onto these affine spaces and approximating each projection using an RVFL network on ℝ^d. To make this more precise, observe that, since each affine projection acts on $x \in M$ as P_j,kx = c_j,k+Φ_j,k(x−c_j,k) for some othogonal projection $Φ_{j, k} : ℝ^{N} \to ℝ^{N}$ , for each k ∈ {1, …K_j}, we have

f (P_{j, k} x) = f (c_{j, k} + Φ_{j, k} (x - c_{j, k})) = f ((I_{N} - Φ_{j, k}) c_{j, k} + U_{j, k} D_{j, k} V_{j, k}^{T} x),

where $Φ_{j, k} = U_{j, k} D_{j, k} V_{j, k}^{T}$ is the compact singular value decomposition (SVD) of Φ_j,k (i.e., only the left and right singular vectors corresponding to non-zero singular values are computed). In particular, the matrix of right-singular vectors $V_{j, k} : ℝ^{d} \to ℝ^{N}$ enables us to define a function ${\hat{f}}_{j, k} : ℝ^{d} \to ℝ$ , given by

\begin{array}{l} {\hat{f}}_{j, k} (z) : = f ((I_{N} - Φ_{j, k}) c_{j, k} + U_{j, k} D_{j, k} z), z \in ℝ^{d}, \end{array}

which satisfies ${\hat{f}}_{j, k} (V_{j, k}^{T} x) = f (P_{j, k} x)$ for all $x \in M$ . By continuity of f and Equation 21, this means that for any ε > 0, there exists j ∈ ℕ such that $| f (x) - {\hat{f}}_{j, k} (V_{j, k}^{T} x) | < ε$ for some k ∈ {1, …, K_j}. For such k ∈ {1, …, K_j}, we may therefore approximate f on the affine space associated with (c_j,k, P_j,k) by approximating ${\hat{f}}_{j, k}$ using a RFVL network ${\tilde{f}}_{n_{j, k}} : ℝ^{d} \to ℝ$ of the form

\begin{array}{l} {\tilde{f}}_{n_{j, k}} (z) : = \sum_{ℓ = 1}^{n_{j, k}} v_{ℓ}^{(j, k)} ρ (〈 w_{ℓ}^{(j, k)}, z 〉 + b_{ℓ}^{(j, k)}), & (22) \end{array}

where ${w_{ℓ}^{(j, k)}}_{ℓ = 1}^{n_{j, k}} \subset ℝ^{d}$ and ${b_{ℓ}^{(j, k)}}_{ℓ = 1}^{n_{j, k}} \subset ℝ$ are random input-to-hidden layer weights and biases (resp.) and the hidden-to-output layer weights ${v_{ℓ}^{(j, k)}}_{ℓ = 1}^{n_{j, k}} \subset ℝ$ are learned. Choosing the activation function ρ and random input-to-hidden layer weights and biases as in Theorem 8 then guarantees that $| f (P_{j, k} x) - {\tilde{f}}_{n_{j, k}} (V_{j, k}^{T} x) |$ is small with high probability whenever n_j,k is sufficiently large.

In light of the above discussion, we propose the following RVFL network construction for approximating functions $f \in C (M)$ : Given a GMRA approximation of $M$ with sufficiently high resolution j, construct and train RVFL networks of the form given by Equation 22 for each k ∈ {1, …, K_j}. Then, given $x \in M$ and ε > 0, choose $k^{'} \in {1, \dots, K_{j}}$ such that

c_{j, k^{'}} \in \underset{c_{j, k} \in C_{j}}{arg min} ∥ x - c_{j, k} ∥_{2}

and evaluate ${\tilde{f}}_{n_{j, k^{'}}} (x)$ to approximate f(x). We summarize this algorithm in Algorithm 1. Since the structure of the GMRA approximation implies $∥ x - P_{j, k^{'}} x ∥_{2} \leq C_{x} 2^{- 2 j}$ holds for our choice of $k^{'} \in {1, \dots, K_{j}}$ [see 43], continuity of f and Lemma 5 imply that, for any ε > 0 and j large enough,

\begin{array}{l} | f (x) - {\tilde{f}}_{n_{j, k^{'}}} (V_{j, k^{'}}^{T} x) | \leq | f (x) - {\hat{f}}_{j, k^{'}} (V_{j, k^{'}}^{T} x) | + | {\hat{f}}_{j, k^{'}} (V_{j, k^{'}}^{T} x) \\ - {\tilde{f}}_{n_{j, k^{'}}} (V_{j, k^{'}}^{T} x) | < ε \end{array}

Algorithm 1

Algorithm 1. Approximation algorithm.

holds with high probability, provided $n_{j, k^{'}}$ satisfies the requirements of Theorem 8.

Remark 7 In the RVFL network construction proposed above, we require that the function f be defined in a sufficiently large region around the manifold. Essentially, we need to ensure that f is continuously defined on the set $S : = M \cup {\hat{M}}_{j}$ , where ${\hat{M}}_{j}$ is the scale-j GMRA approximation

{\hat{M}}_{j} : = {P_{j, k_{j} (z)} z : || z | |_{2} \leq rad (M)} \cap B_{2}^{N} (0, rad (M)) .

This ensures that f can be evaluated on the affine subspaces given by the GMRA.

To simulate Algorithm 1, we take $M = S^{2}$ embedded in ℝ²⁰ and construct a GMRA up to level j_max = 15 using 20,000 data points sampled uniformly from $M$ . Given j ≤ j_max, we generate RVFL networks ${\hat{f}}_{n_{j, k}} : ℝ^{2} \to ℝ$ as in Equation 22 and train them on $V_{j, k}^{T} (B_{2}^{N} (c_{j, k}, r) \cap T_{j, k})$ using the training pairs ${(V_{k, j}^{T} x_{ℓ}, f (P_{j, k} x_{ℓ}))}_{ℓ = 1}^{p}$ , where T_k,j is the affine space generated by (c_j,k, P_j,k). For simplicity, we fix n_j,k = n to be constant for all k ∈ {1, …, K_j} and use a single, fixed pair of parameters α, Ω > 0 when constructing all RVFL networks. We then randomly select a test set of 200 points $x \in M$ for use throughout all experiments. In each experiment (i.e., point in Figure 1), we use Algorithm 1 to produce an approximation $y^{♯} = {\tilde{f}}_{n_{j, k^{'}}} (x)$ of f(x). Figure 1 displays the mean relative error in these approximations for varying numbers of nodes n; to construct this plot, f is taken to be the exponential $f (x) = exp (\sum_{k = 1}^{N} x (k))$ and ρ the hyperbolic secant function. Notice that for small numbers of nodes, the RVFL networks are not very good at approximating f, regardless of the choice of α, Ω > 0. However, the error decays as the number of nodes increases until reaching a floor due to error inherent in the GMRA approximation. Hence, as suggested by Theorem 3, to achieve a desired error bound of ε > 0, one needs to only choose a GMRA scale j such that the inherent error in the GMRA (which scales like 2^−j) is less than ε, then adjust the parameters α_j, Ω_j, and n_j,k accordingly.

Figure 1

Figure 1. Log-scale plot of average relative error for Algorithm 1 as a function of the number of nodes n in each RVFL network. Black (cross), blue (circle), and red (square) lines correspond to GMRA refinement levels j = 12, j = 9, and j = 6 (resp.). For each j, we fix α_j = 2 and vary Ω_j = 10, 15 (solid and dashed lines, resp.). Reconstruction error decays as a function of n until reaching a floor due to error in the GMRA approximation of $M$ . The code used to obtain these numerical results is available upon direct request sent to the corresponding author.

Remark 8 As we just mentioned, the error can only decay so far due to the resolution of the GMRA approximation. However, that is not the only floor in our simulation; indeed, the ε in Theorem 3 is determined by the α_j's and Ω_j's, which we kept fixed (see the caption of Figure 1). Consequently, the stagnating accuracy as n increases, as seen in Figure 1, is also predicted by Theorem 3. Since the solid and dashed lines seem to reach the same floor, the floor due to error inherent in the GMRA approximation seems to be the limiting error term for RVFL networks with large numbers of nodes.

Remark 9 Utilizing random inner weights and biases resulted in us needing to approximate the atlas to the manifold. To this end, knowing the computational complexity of the GMRA approximation would be useful in practice. As it turns out in Liao and Maggioni [45], calculating the GMRA approximation has computational complexity O(C^dNmlog(m)), where m is the number of training data points and C > 0 is a numerical constant.

3.5 Proofs of technical lemmas

3.5.1 Proof of Lemma 3

Observe that h_Ω defined in Equation 3 may be viewed as a multidimensional bump function; indeed, the parameter Ω > 0 controls the width of the bump. In particular, if Ω is allowed to grow very large, then h_Ω becomes very localized near the origin. Objects that behave in this way are known in the functional analysis literature as approximate δ-functions:

Definition 2 A sequence of functions ${φ_{t}}_{t > 0} \subset L^{1} (ℝ^{N})$ are called approximate (or nascent) δ-functions if

lim_{t \to \infty} \int_{ℝ^{N}} φ_{t} (x) f (x) d x = f (0)

for all $f \in C_{c} (ℝ^{N})$ . For such functions, we write $δ_{0} (x) = lim_{t \to \infty} φ_{t} (x)$ for all x ∈ ℝ^N, where δ₀ denotes the N-dimensional Dirac δ-function centered at the origin.

Given φ ∈ L¹(ℝ^N) with $\int_{ℝ^{N}} φ (x) d x = 1$ , one may construct approximate δ-functions for t > 0 by defining $φ_{t} (x) : = t^{N} φ (t x)$ for all x ∈ ℝ^N [46]. Such sequences of approximate δ-functions are also called approximate identity sequences [47] since they satisfy a particularly nice identity with respect to convolution, namely, $lim_{t \to \infty} ∥ f * φ_{t} - f ∥_{1} = 0$ for all $f \in C_{c} (ℝ^{N})$ [see 47, Theorem 6.32]. In fact, such an identity holds much more generally.

Lemma 6 [46, Theorem 1.18] Let φ ∈ L¹(ℝ^N) with $\int_{ℝ^{N}} φ (x) d x = 1$ and for t > 0 define $φ_{t} (x) : = t^{N} φ (t x)$ for all x ∈ ℝ^N. If f ∈ L^p(ℝ^N) for 1 ≤ p < ∞ (or $f \in C_{0} (ℝ^{N}) \subset L^{\infty} (ℝ^{N})$ for p = ∞), then $lim_{t \to \infty} ∥ f * φ_{t} - f ∥_{p} = 0$ .

To prove Equation 4, it would suffice to have $lim_{Ω \to \infty} ∥ f * h_{Ω} - f ∥_{\infty} = 0,$ which is really just Lemma 6 in case p = ∞. Nonetheless, we present a proof by mimicking [46] for completeness. Moreover, we will use a part of proof in Remark 10 below.

Lemma 7 Let h ∈ L¹(ℝⁿ) with $\int_{ℝ^{N}} h (x) d x = 1$ and define $h_{Ω} \in L^{1} (ℝ^{N})$ as in Equation 3 for all Ω > 0. Then, for all $f \in C_{0} (ℝ^{N})$ , we have

lim_{Ω \to \infty} sup_{x \in ℝ^{N}} | (f * h_{Ω}) (x) - f (x) | = 0 .

Proof. By symmetry of the convolution operator in its arguments, we have

\begin{array}{l} \sup_{x \in ℝ^{N}} | (f * h_{Ω}) (x) - f (x) | = \sup_{x \in ℝ^{N}} | \int_{ℝ^{N}} f (y) h_{Ω} (x - y) d y - f (x) | \\ = \sup_{x \in ℝ^{N}} | \int_{ℝ^{N}} f (x - y) h_{Ω} (y) d y - f (x) | . \end{array}

Since a simple substitution yields $1 = \int_{ℝ^{N}} h (x) d x = \int_{ℝ^{N}} h_{Ω} (x) d x,$ it follows that

\begin{array}{r} sup_{x \in ℝ^{N}} | (f * h_{Ω}) (x) - f (x) | = sup_{x \in ℝ^{N}} | \int_{ℝ^{N}} | f (x - y) - f (x) | h_{Ω} (y) d y | \\ \leq \int_{ℝ^{N}} | h_{Ω} (y) | sup_{x \in ℝ^{N}} | f (x) - f (x - y) | d y . \end{array}

Finally, expanding the function h_Ω, we obtain

\begin{array}{l} sup_{x \in ℝ^{N}} | (f * h_{Ω}) (x) - f (x) | \\ \leq \int_{ℝ^{N}} (Ω^{N} | h (Ω y) |) sup_{x \in ℝ^{N}} | f (x) - f (x - y) | d y \\ = \int_{ℝ^{N}} | h (z) | sup_{x \in ℝ^{N}} | f (x) - f (x - z / Ω) | d z, \end{array}

where we have used the substitution z = Ωy. Taking limits on both sides of this expression and observing that

\int_{ℝ^{N}} | h (z) | sup_{x \in ℝ^{N}} | f (x) - f (x - z / Ω) | d z \leq 2 ∥ h ∥_{1} sup_{x \in ℝ^{N}} | f (x) | < \infty,

using the Dominated Convergence Theorem, we obtain

\begin{array}{l} lim_{Ω \to \infty} sup_{x \in ℝ^{N}} | (f * h_{Ω}) (x) - f (x) | \\ \leq \int_{ℝ^{N}} | h (z) | {lim}_{Ω \to \infty} {sup}_{x \in ℝ^{N}} | f (x) - f (x - z / Ω) | d z . \end{array}

So, it suffices to show that, for all z ∈ ℝ^N,

lim_{Ω \to \infty} sup_{x \in ℝ^{N}} | f (x) - f (x - z / Ω) | = 0 .

To this end, let ε > 0 and z ∈ ℝ^N be arbitrary. Since $f \in C_{0} (ℝ^{N})$ , there exists r > 0 sufficiently large such that |f(x)| < ε/2 for all $x \in ℝ^{N} \ \bar{B (0, r)}$ , where $\bar{B (0, r)} \subset ℝ^{N}$ is the closed ball of radius r centered at the origin. Let $B : = \bar{B (0, r + ∥ z / Ω ∥_{2})},$ so that for each $x \in ℝ^{N} \ B$ we have both x and x−z/Ω in $ℝ^{N} \ \bar{B (0, r)}$ . Thus, both |f(x)| < ε/2 and |f(x−z/Ω)| < ε/2, implying that

sup_{x \in ℝ^{N} \ B} | f (x) - f (x - z / Ω) | < ε .

Hence, we obtain

\begin{array}{l} lim_{Ω \to \infty} sup_{x \in ℝ^{N}} | f (x) - f (x - z / Ω) | \\ \leq lim_{Ω \to \infty} max {sup_{x \in B} | f (x) - f (x - z / Ω) |, \\ sup_{x \in ℝ^{N} \ B} | f (x) - f (x - z / Ω) |} \\ \leq max {ε, lim_{Ω \to \infty} sup_{x \in B} | f (x) - f (x - z / Ω) |} . \end{array}

Now, as $B$ is a compact subset of ℝ^N, the continuous function f is uniformly continuous on $B$ , and so the remaining limit and supremum may be freely interchanged, whereby continuity of f yields

lim_{Ω \to \infty} sup_{x \in B} | f (x) - f (x - z / Ω) | = sup_{x \in B} lim_{Ω \to \infty} | f (x) - f (x - z / Ω) | = 0 .

Since ε > 0 may be taken arbitrarily small, we have proved the result.

Remark 10 While Lemma 7 does the approximation we aim for, it gives no indication of how fast

sup_{x \in ℝ^{N}} | (f * h_{Ω}) (x) - f (x) |

decays in terms of Ω or the dimension N. Assuming h(z) = g(z(1))⋯g(z(N)) for some non-negative g (which is how we will choose h in Section 3.5.2) and f to be β-Hölder continuous for some fixed β ∈ (0, 1) yields that

\begin{array}{l} sup_{x \in ℝ^{N}} | (f * h_{Ω}) (x) - f (x) | \\ \leq \int_{ℝ^{N}} | h (z) | sup_{x \in ℝ^{N}} | f (x) - f (x - z / Ω) | d z \\ ≲ Ω^{- β} \int_{ℝ^{N}} ∥ z ∥_{2}^{β} h (z) d z \\ \leq Ω^{- β} (\int_{ℝ^{N}} (z {(1)}^{2} + \dots + z {(N)}^{2}) h (z) d z)^{β / 2} \\ \leq Ω^{- β} (N max_{j \in {1, \dots, N}} \int_{ℝ^{N}} z {(j)}^{2} h (z) d z)^{β / 2} \\ = {(\sqrt{N} / Ω)}^{β} (max_{j \in {1, \dots, N}} \int_{ℝ} z {(j)}^{2} g (z (j)) d z (j))^{β / 2} \\ = {(\sqrt{N} / Ω)}^{β} (\int_{ℝ} z {(1)}^{2} g (z (1)) d z (1))^{β / 2} \\ ≲ {(\sqrt{N} / Ω)}^{β} \end{array}

where the third inequality follows from Jensen's inequality.

3.5.2 Proof of Lemma 4: the limit-integral representation

Let A ∈ C^∞(ℝ) be any even function supported on $[- \frac{1}{2}, \frac{1}{2}]$ s.t. ∥A∥₂ = 1. Then, ϕ = A*A ∈ C^∞(ℝ) is an even function supported on [−1, 1] s.t. ϕ(0) = 1. Lemma 3 implies that

\begin{array}{l} f (x) = lim_{Ω \to \infty} (f * h_{Ω}) (x) & (23) \end{array}

uniformly in x ∈ K for any h ∈ L¹(ℝ^N) satisfying $\int_{ℝ^{N}} h (z) d z = 1 .$ We choose

h (z) = \frac{1}{{(2 π)}^{N}} \int_{ℝ^{N}} exp (i 〈 w, z 〉) \prod_{j = 1}^{N} ϕ (w (j)) d w

which the reader may recognize as the (inverse) Fourier transform of $\prod_{j = 1}^{N} ϕ (w (j))$ . As we announced in Remark 10, h(z) = g(z(1))⋯g(z(N)), where (using the convolution theorem)

\begin{array}{l} g (z (j)) = \frac{1}{2 π} \int_{ℝ} exp (i w (j) z (j)) ϕ (w (j)) d w (j) \\ = \frac{1}{2 π} \int_{ℝ} exp (i w (j) z (j)) (A * A) (w (j)) d w (j) \\ = 2 π (\frac{1}{2 π} \int_{ℝ} exp (i w (j) z (j)) A (w (j)) d w (j))^{2} \geq 0 \end{array}

Moreover, since g is the Fourier transform of an even function, h is real-valued and also even. In addition, since ϕ is smooth, h decays faster than the reciprocal of any polynomial (as follows from repeated integration by parts and the Riemann–Lebesgue lemma), so h ∈ L¹(ℝ^N). Thus, Fourier inversion yields

\int_{ℝ^{N}} h (z) d z = \int_{ℝ^{N}} exp (- i 〈 w, z 〉) h (z) d z |_{w = 0} = \prod_{j = 1}^{N} ϕ (0) = 1,

which justifies our application of Lemma 3. Expanding the right-hand side of Equation 23 (using the scaling property of the Fourier transform) yields that

\begin{array}{l} (f * h_{Ω}) (x) = \int_{ℝ^{N}} f (y) h_{Ω} (x - y) d y \\ = \frac{1}{{(2 π)}^{N}} \int_{K} f (y) \int_{ℝ^{N}} exp (i 〈 w, x - y 〉) \\ \prod_{j = 1}^{N} ϕ (w (j) / Ω) d w d y \\ = \frac{1}{{(2 π)}^{N}} \int_{K} \int_{{[- Ω, Ω]}^{N}} f (y) cos (〈 w, x - y 〉) \\ \prod_{j = 1}^{N} ϕ (w (j) / Ω) d w d y & (24) \end{array}

because ϕ is even and supported on [−1, 1]. Since Equation 24 is an iterated integral of a continuous function over a compact set, Fubini's theorem readily applies, yielding

\begin{array}{l} f (x) = lim_{Ω \to \infty} (f * h_{Ω}) (x) \\ = lim_{Ω \to \infty} \frac{1}{{(2 π)}^{N}} \int_{K \times {[- Ω, Ω]}^{N}} f (y) cos (〈 w, x - y 〉) \prod_{j = 1}^{N} ϕ (w (j) / Ω) d y d w . \end{array}

Since $| 〈 w, x - y 〉 | \leq ∥ x - y ∥_{1} ∥ w ∥_{\infty} \leq 2 N rad (K) Ω \leq (L + \frac{1}{2}) π,$ it follows that

\begin{array}{r} f (x) = lim_{Ω \to \infty} \frac{1}{{(2 π)}^{N}} \int_{K \times {[- Ω, Ω]}^{N}} f (y) {cos}_{Ω} (〈 w, x - y 〉) \\ \prod_{j = 1}^{N} ϕ (w (j) / Ω) d y d w & (25) \end{array}

where cos_Ω is defined in Equation 5.

With the representation given by Equation 25 in hand, we now seek to reintroduce the general activation function ρ. To this end, since cos_Ω ∈ C_c(ℝ)⊂C₀(ℝ) we may apply the convolution identity given by Equation 4 with f replaced by cos_Ω to obtain ${cos}_{Ω} (z) = lim_{α \to \infty} ({cos}_{Ω} * h_{α}) (z)$ uniformly for all z ∈ ℝ, where h_α(z) = αρ(αz). Using this representation of cos_Ω in Equation 25, it follows that

\begin{array}{r} f (x) = lim_{Ω \to \infty} \frac{1}{{(2 π)}^{N}} \int_{K \times {[- Ω, Ω]}^{N}} f (y) (lim_{α \to \infty} ({cos}_{Ω} * h_{α}) (〈 w, x - y 〉)) \\ \prod_{j = 1}^{N} ϕ (w (j) / Ω) d y d w \end{array}

holds uniformly for all x ∈ K. Since f is continuous and the convolution cos_Ω*h_α is uniformly continuous and uniformly bounded in α by ∥ρ∥₁ (see below), the fact that the domain K×[−Ω, Ω]^N is compact then allows us to bring the limit as α tends to infinity outside the integral in this expression via the Dominated Convergence Theorem, which gives us

\begin{array}{r} f (x) = lim_{Ω \to \infty} lim_{α \to \infty} \frac{1}{{(2 π)}^{N}} \int_{K \times {[- Ω, Ω]}^{N}} f (y) ({cos}_{Ω} * h_{α}) \\ (〈 w, x - y 〉) \prod_{j = 1}^{N} ϕ (w (j) / Ω) d y d w & (26) \end{array}

uniformly for every x ∈ K. The uniform boundedness of the convolution follows from the fact that

\begin{array}{l} ({cos}_{Ω} * h_{α}) (z) = \int_{ℝ} {cos}_{Ω} (z - u) h_{α} (u) d u \\ = \int_{ℝ} {cos}_{Ω} (z - v / α) ρ (v) d v, & (27) \end{array}

where v = αu.

Remark 11 It should be noted that we are unable to swap the order of the limits in Equation 26, since cos_Ω is not in C₀(ℝ) when Ω is allowed to be infinite.

Remark 12 Complementing Remark 10, we will now elucidate how fast

\begin{array}{l} | {cos}_{Ω} (z) - ({cos}_{Ω} * h_{α}) (z) | \end{array}

decays in terms of α. Using the fact that $\int_{ℝ} ρ (z) d z = 1,$ Equation 27 and the triangle inequality allows us to bound the absolute difference above by

\begin{array}{l} \int_{ℝ} | {cos}_{Ω} (z) - {cos}_{Ω} (z - v / α) | \cdot | ρ (v) | d v . \end{array}

Since cos_Ω is 1-Lipschitz, it follows that the above integral is bounded by $\int_{ℝ} | v ρ (v) | d v / α .$

To complete this step of the proof, observe that the definition of cos_Ω allows us to write

\begin{array}{l} ({cos}_{Ω} * h_{α}) (z) = α \int_{ℝ} {cos}_{Ω} (u) ρ | α (z - u) | d u \\ = α \int_{- \frac{π}{2} (2 L + 1)}^{\frac{π}{2} (2 L + 1)} {cos}_{Ω} (u) ρ (α (z - u)) d u & (28) \end{array}

By substituting Equation 28 into Equation 26, we then obtain

\begin{array}{r} f (x) = lim_{Ω \to \infty} lim_{α \to \infty} \frac{α}{{(2 π)}^{N}} \int_{K (Ω)} f (y) {cos}_{Ω} (u) ρ (α (〈 w, x - y 〉 - u)) \\ \prod_{j = 1}^{N} ϕ (w (j) / Ω) d y d w d u \end{array}

uniformly for all x ∈ K, where $K (Ω) : = K \times {[- Ω, Ω]}^{N} \times [- \frac{π}{2} (2 L + 1), \frac{π}{2} (2 L + 1)]$ . In this way, recalling that $F_{α, Ω} (y, w, u) : = \frac{α}{{(2 π)}^{N}} f (y) {cos}_{Ω} (u) \prod_{j = 1}^{N} ϕ (w (j) / Ω),$ and b_α(y, w, u): = −α(〈w, y〉+u) for y, w ∈ ℝ^N and u ∈ ℝ, we conclude the proof.

3.5.3 Proof of Lemma 5: Monte-Carlo integral approximation

The next step in the proof of Theorem 5 is to approximate the integral in Equation 7 using the Monte-Carlo method. To this end, let ${y_{k}}_{k = 1}^{n}$ , ${w_{k}}_{k = 1}^{n}$ , and ${u_{k}}_{k = 1}^{n}$ be independent samples drawn uniformly from K, [−Ω, Ω]^N, and $[- \frac{π}{2} (2 L + 1), \frac{π}{2} (2 L + 1)]$ , respectively, and consider the sequence of random variables ${I_{n} (x)}_{n = 1}^{\infty}$ defined by

\begin{array}{l} I_{n} (x) : = \frac{vol (K (Ω))}{n} \sum_{k = 1}^{n} F_{α, Ω} (y_{k}, w_{k}, u_{k}) ρ (α 〈 w_{k}, x 〉 + b_{α} (y_{k}, w_{k}, u_{k})) & (29) \end{array}

for each x ∈ K, where we note that vol(K(Ω)) = (2Ω)^Nπ(2L + 1)vol(K). If we also define

\begin{array}{l} I (x; p) : = \int_{K (Ω)} (F_{α, Ω} (y, w, u) ρ (α 〈 w, x 〉 + b_{α} (y, w, u)))^{p} d y d w d u \end{array}

for x ∈ K and p ∈ ℕ, then we want to show that

\begin{array}{l} 𝔼 \int_{K} | I (x; 1) - I_{n} (x) |^{2} d x = O (1 / n) & (30) \end{array}

as n → ∞, where the expectation is taken with respect to the joint distribution of the random samples ${y_{k}}_{k = 1}^{n}$ , ${w_{k}}_{k = 1}^{n}$ , and ${u_{k}}_{k = 1}^{n}$ . For this, it suffices to find a constant C(f, ρ, α, Ω, N) < ∞ independent of n satisfying

\int_{K} 𝔼 | I (x; 1) - I_{n} (x) |^{2} d x \leq \frac{C (f, ρ, α, Ω, N)}{n} .

Indeed, an application of Fubini's theorem would then yield

𝔼 \int_{K} | I (x; 1) - I_{n} (x) |^{2} d x \leq \frac{C (f, ρ, α, Ω, N)}{n},

which implies Equation 30. To determine such a constant, we first observe by Theorem 4 that

𝔼 | I (x; 1) - I_{n} (x) |^{2} = \frac{vo l^{2} (K (Ω)) σ {(x)}^{2}}{n},

where we define the variance term

σ {(x)}^{2} : = \frac{I (x; 2)}{vol (K (Ω))} - \frac{I {(x; 1)}^{2}}{vo l^{2} (K (Ω))}

for x ∈ K. Since ∥ϕ∥_∞ = 1 (see Lemma 8 below), it follows that

| F_{α, Ω} (y, w, u) | = \frac{α}{{(2 π)}^{N}} | f (y) | \cdot | {cos}_{Ω} (u) | \prod_{j = 1}^{N} | ϕ (w (j) / Ω) | \leq \frac{α M}{{(2 π)}^{N}}

for all y, w ∈ ℝ^N and u ∈ ℝ, where $M : = sup_{x \in K} | f (x) | < \infty$ , we obtain the following simple bound on the variance term

\begin{array}{l} σ {(x)}^{2} \leq \frac{I (x; 2)}{vol (K (Ω))} \leq \frac{α^{2} M^{2}}{{(2 π)}^{2 N} vol (K (Ω))} \\ \int_{K (Ω)} | ρ (α 〈 w, x 〉 + b_{α} (y, w, u)) |^{2} d y d w d u . \end{array}

Since we assume ρ ∈ L^∞(ℝ), we then have

\begin{array}{c} \int_{K} 𝔼 | I (x; 1) - I_{n} (x) |^{2} d x = \frac{vo l^{2} (K (Ω))}{n} \int_{K} σ {(x)}^{2} d x \\ \leq \frac{α^{2} M^{2} vol (K (Ω))}{{(2 π)}^{2 N} n} \int_{K \times K (Ω)} | ρ (α 〈 w, x 〉 + b_{α} (y, w, u)) |^{2} d x d y d w d u \\ = \frac{α^{2} M^{2} vo l^{2} (K (Ω)) vol (K) ‖ ρ ‖_{\infty}^{2}}{{(2 π)}^{2 N} n} . \end{array}

Substituting the value of vol(K(Ω)), we obtain

C (f, ρ, α, Ω, N) : = α^{2} M^{2} {(Ω / π)}^{2 N} π^{2} {(2 L + 1)}^{2} vo l^{3} (K) ‖ ρ ‖_{\infty}^{2}

is a suitable choice for the desired constant.

Now that we have established Equation 30, we may rewrite the random variables I_n(x) in a more convenient form. To this end, we change the domain of the random samples ${w_{k}}_{k = 1}^{n}$ to [−αΩ, αΩ]^N and define the new random variables ${b_{k}}_{k = 1}^{n} \subset ℝ$ by b_k: = −(〈w_k, y_k〉+αu_k) for each k = 1, …, n. In this way, if we denote

v_{k} : = \frac{vol (K (Ω))}{n} F_{α, Ω} (y_{k}, \frac{w_{k}}{α}, u_{k})

for each k = 1, …, n, the random variables ${f_{n}}_{n = 1}^{\infty}$ defined by

f_{n} (x) : = \sum_{k = 1}^{n} v_{k} ρ (〈 w_{k}, x 〉 + b_{k})

satisfy f_n(x) = I_n(x) for every x ∈ K. Combining this with Equation 30, we have proved Lemma 5.

Lemma 8 ∥ϕ∥_∞ = 1.

Proof. It suffices to prove that |ϕ(z)| ≤ 1 for all z ∈ ℝ because ϕ(0) = 1. By Cauchy–Schwarz,

\begin{array}{l} | ϕ (z) | = | \int_{ℝ} A (u) A (z - u) d u | \\ \leq \sqrt{\int_{ℝ} A (u) A (u) d u \int_{ℝ} A (z - u) A (z - u) d u} \\ = \sqrt{\int_{ℝ} A (u) A (0 - u) d u \int_{ℝ} A (v) A (- v) d v} = \sqrt{ϕ (0) ϕ (0)} = 1 \end{array}

because A is even.

3.5.4 Proof of Theorem 1 when ρ′ ∈ L¹(ℝ)∩L^∞(ℝ)

Let $f \in C_{c} (ℝ^{N})$ with K: = supp(f) and suppose ε > 0 is fixed. Take the activation function ρ:ℝ → ℝ to be differentiable with ρ′ ∈ L¹(ℝ)∩L^∞(ℝ). We wish to show that there exists a sequence of RVFL networks ${f_{n}}_{n = 1}^{\infty}$ defined on K which satisfy the asymptotic error bound

𝔼 \int_{K} | f (x) - f_{n} (x) |^{2} d x \leq ε + O (1 / n)

as n → ∞. The proof of this result is a minor modification of second step in the proof of Theorem 5.

If we redefine h_α(z) as αρ′(αz), then Equation 26 plainly still holds and Equation 28 reads

({cos}_{Ω} * h_{α}) (z) = α \int_{ℝ} {cos}_{Ω} (u) ρ^{'} (α (z - u)) d u .

Recalling the definition of cos_Ω in Equation 5 and integrating by parts, we obtain

\begin{array}{l} ({cos}_{Ω} * h_{α}) (z) = α \int_{ℝ} {cos}_{Ω} (u) ρ^{'} (α (z - u)) d u \\ = - \int_{- \frac{π}{2} (2 L + 1)}^{\frac{π}{2} (2 L + 1)} {cos}_{Ω} (u) d ρ (α (z - u)) \\ = - {cos}_{Ω} (u) ρ (α (z - u)) |_{- \frac{π}{2} (2 L + 1)}^{\frac{π}{2} (2 L + 1)} \\ + \int_{- \frac{π}{2} (2 L + 1)}^{\frac{π}{2} (2 L + 1)} ρ (α (z - u)) d {cos}_{Ω} (u) \\ = - \int_{ℝ} {sin}_{Ω} (u) ρ (α (z - u)) d u \end{array}

for all z ∈ ℝ, where $L : = ⌈ \frac{2 N}{π} rad (K) Ω - \frac{1}{2} ⌉$ and sin_Ω:ℝ → [−1, 1] is defined analogously to Equation 5. Substituting this representation of (cos_Ω*h_α)(z) into Equation 26 then yields

\begin{array}{l} f (x) = lim_{Ω \to \infty} lim_{α \to \infty} \frac{- α}{{(2 π)}^{N}} \int_{K (Ω)} f (y) {sin}_{Ω} (u) ρ (α (〈 w, x - y 〉 - u)) \\ \prod_{j = 1}^{N} ϕ (w (j) / Ω) d y d w d u \end{array}

uniformly for every x ∈ K. Thus, if we replace the definition of F_{α, Ω} in Equation 6 by

F_{α, Ω} (y, w, u) : = \frac{- α}{{(2 π)}^{N}} f (y) {sin}_{Ω} (u) \prod_{j = 1}^{N} ϕ (w (j) / Ω)

for y, w ∈ ℝ^N and u ∈ ℝ, we again obtain the uniform representation given by Equation 7 for all x ∈ K. The remainder of the proof proceeds from this point exactly as in the proof of Theorem 5.

3.5.5 Proof of Theorem 8

We wish to show that there exist sequences of RVFL networks ${{\tilde{f}}_{n_{j}}}_{n_{j} = 1}^{\infty}$ defined on ϕ_j(U_j) for each j ∈ J, which together satisfy the error bound

\int_{M} | f (x) - \sum_{{j \in J : x \in U_{j}}} ({\tilde{f}}_{n_{j}} \circ ϕ_{j}) (x) |^{2} d x < ε

with probability at least 1−η for {_{n_j}j ∈ J} sufficiently large. The proof is obtained by showing that

\begin{array}{l} | f (x) - \sum_{{j \in J : x \in U_{j}}} ({\tilde{f}}_{n_{j}} \circ ϕ_{j}) (x) | < \sqrt{\frac{ε}{vol (M)}} & (31) \end{array}

holds uniformly for $x \in M$ with high probability.

We begin as in the proof of Theorem 7 by applying the representation given by Equation 17 for (f on each chart (U_j, ϕ_j), which gives us

\begin{array}{l} | f (x) - \sum_{{j \in J : x \in U_{j}}} ({\tilde{f}}_{n_{j}} \circ ϕ_{j}) (x) | \\ \leq \sum_{{j \in J : x \in U_{j}}} | ({\hat{f}}_{j} \circ ϕ_{j}) (x) - ({\tilde{f}}_{n_{j}} \circ ϕ_{j}) (x) | & (32) \end{array}

for all $x \in M$ . Now, since we have already seen that ${\hat{f}}_{j} \in C_{c} (ℝ^{d})$ for each j ∈ J, Theorem 6 implies that for any ε_j > 0, there exist constants α_j, Ω_j > 0 and hidden-to-output layer weights ${v_{k}^{(j)}}_{k = 1}^{n_{j}} \subset ℝ$ for each j ∈ J such that for any

\begin{array}{l} δ_{j} < \frac{\sqrt{ε_{j}}}{8 \sqrt{2 d} κ α_{j}^{2} M_{j} Ω_{j} {(Ω_{j} / π)}^{d} vo l^{3 / 2} (ϕ_{j} (U_{j})) (π + 2 d rad (ϕ_{j} (U_{j})) Ω)} & (33) \end{array}

we have

| {\hat{f}}_{j} (z) - {\tilde{f}}_{n_{j}} (z) | < \sqrt{\frac{ε_{j}}{2 vol (ϕ_{j} (U_{j}))}}

uniformly for all z ∈ ϕ_j(U_j) with probability at least 1−η_j, provided the number of nodes n_j satisfies

\begin{array}{l} n_{j} & \geq \frac{c Σ^{(j)} α_{j} {(Ω_{j} / π)}^{d} (π + 2 d rad (ϕ_{j} (U_{j})) Ω_{j}) log (3 η_{j}^{- 1} N (δ_{j}, ϕ_{j} (U_{j})))}{\sqrt{ε_{j}} log (1 + \frac{\sqrt{ε_{j}}}{Σ^{(j)} α_{j} {(Ω_{j} / π)}^{d} (π + 2 d rad (ϕ_{j} (U_{j})) Ω_{j})})}, & (34) \end{array}

where c > 0 is a numerical constant and $Σ^{(j)} : = 2 C^{(j)} \sqrt{2 vol (ϕ_{j} (U_{j}))} .$ Indeed, it suffices to choose

v_{k}^{(j)} : = \frac{vol (K (Ω_{j}))}{n_{j}} F_{α_{j}, Ω_{j}} (y_{k}^{(j)}, \frac{w_{k}^{(j)}}{α_{j}}, u_{k}^{(j)})

for each k = 1, …, n_j, where

K (Ω_{j}) : = ϕ_{j} (U_{j}) \times {[- α_{j} Ω_{j}, α_{j} Ω_{j}]}^{d} \times [- \frac{π}{2} (2 L_{j} + 1), \frac{π}{2} (2 L_{j} + 1)]

for each j ∈ J. Combined with Equation 32, choosing δ_j and n_j satisfying Equations 33, 34, respectively, then yields

\begin{array}{l} | f (x) - \sum_{{j \in J : x \in U_{j}}} ({\tilde{f}}_{n_{j}} \circ ϕ_{j}) (x) | < \sum_{{j \in J : x \in U_{j}}} \sqrt{\frac{ε_{j}}{2 vol (ϕ_{j} (U_{j}))}} \\ \leq \sum_{j \in J} \sqrt{\frac{ε_{j}}{2 vol (ϕ_{j} (U_{j}))}} \end{array}

for all $x \in M$ with probability at least $1 - \sum_{{j \in J : x \in U_{j}}} η_{j} \geq 1 - \sum_{j \in J} η_{j}$ . Since we require that Equation 31 holds for all $x \in M$ with probability at least 1−η, the proof is then completed by choosing {_{ε_j}j ∈ J} and {_{η_j}j ∈ J}, such that

ε = \frac{vol (M)}{2} (\sum_{j \in J} \sqrt{\frac{ε_{j}}{vol (ϕ_{j} (U_{j}))}})^{2} and η = \sum_{j \in J} η_{j} .

In particular, it suffices to choose

ε_{j} = \frac{2 vol (ϕ_{j} (U_{j})) ε}{| J |^{2} vol (M)}

and η_j = η/|J| for each j ∈ J, so that Equations 33, 34 become

\begin{array}{l} δ_{j} & < \frac{\sqrt{ε}}{8 | J | \sqrt{d vol (M)} κ α_{j}^{2} M_{j} Ω_{j} {(Ω_{j} / π)}^{d} vol (ϕ_{j} (U_{j})) (π + 2 d rad (ϕ_{j} (U_{j})) Ω)}, \\ n_{j} & \geq \frac{2 c | J | \sqrt{vol (M)} C^{(j)} α_{j} {(Ω_{j} / π)}^{d} (π + 2 d rad (ϕ_{j} (U_{j})) Ω_{j}) log (3 | J | η^{- 1} N (δ_{j}, ϕ_{j} (U_{j})))}{\sqrt{ε} log (1 + \frac{\sqrt{ε}}{2 | J | \sqrt{vol (M)} C^{(j)} α_{j} {(Ω_{j} / π)}^{d} (π + 2 d rad (ϕ_{j} (U_{j})) Ω_{j})})}, \end{array}

as desired.

4 Discussion

The central topic of this study is the study of the approximation properties of a randomized variation of shallow neural networks known as RVFL. In contrast with the classical single-layer neural networks, training of an RVFL involves only learning the output weights, while the input weights and biases of all the nodes are selected at random from an appropriate distribution and stay fixed throughout the training. The main motivation for studying the properties of such networks is as follows:

1. Random weights are often utilized as an initialization for a NN training procedure. Thus, establishing the properties of the RVFL networks is an important first step toward understanding how random weights are transformed during training.

2. Due to their much more computationally efficient training process, the RVFL networks proved to be a valuable alternative to the classical SLFNs. They were successfully used in several modern applications, especially those that require frequent re-training of a neural network [20, 26, 27].

Despite their practical and theoretical importance, results providing rigorous mathematical analysis of the properties of RVFLs are rare. The work of Igelnik and Pao [16] showed that RVFL networks are universal approximators for the class of continuous, compactly supported functions and established the asymptotic convergence rate of the expected approximation error as a function of the number of nodes in the hidden layer. While this result served as a theoretical justification for using RVFL networks in practice, a close examination led us to the conclusion that the proofs in Igelnik and Pao [16] contained several technical errors.

In this study, we offer a revision and a modification of the proof methods from Igelnik and Pao [16] that allow us to prove a corrected, slightly weaker version of the result announced by Igelnik and Pao. We further build upon their work and show a non-asymptotic probabilistic (instead of on average) approximation result, which gives an explicit bound on the number of hidden layer nodes that are required to achieve the desired approximation accuracy with the desired level of certainty (that is, with high enough probability). In addition to that, we extend the obtained result to the case when the function is supported on a compact, low-dimensional sub-manifold of the ambient space.

While our study closes some of the gaps in the study of the approximation properties of RVFL, we believe that it just starts the discussion and opens many directions for further research. We briefly outline some of them here.

In our results, the dependence of the required number n of the nodes in the hidden layer on the dimension N of the domain is superexponential, which is likely an artifact of the proof methods we use. We believe this dependence can be improved to be exponential by using a different, more refined approach to the construction of the limit-integral representation of a function. A related interesting direction for future research is to study how the bound on n changes for more restricted classes of (e.g., smooth) functions.

Another important direction that we did not discuss in this study is learning the output weights and studying the robustness of the RVFL approximation to the noise in the training data. Obtaining provable robustness guarantees for an RVFL training procedure would be a step toward the robustness analysis of neural networks.

Data availability statement

The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.

Author contributions

DN: Writing – original draft, Writing – review & editing. AN: Writing – original draft, Writing – review & editing. RS: Writing – original draft, Writing – review & editing. PS: Writing – original draft, Writing – review & editing. OS: Writing – original draft, Writing – review & editing.

Funding

The author(s) declare financial support was received for the research, authorship, and/or publication of this article. DN was partially supported by NSF DMS 2108479 and NSF DMS 2011140. RS was partially supported by a UCSD senate research award and a Simons fellowship. PS was partially supported by NSF DMS 1909457.

Acknowledgments

The authors thank F. Krahmer, S. Krause-Solberg, and J. Maly for sharing their GMRA code, which they adapted from that provided by M. Maggioni.

Conflict of interest

The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Publisher's note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Author disclaimer

The views expressed in this article are those of the authors and do not reflect the official policy or position of the U.S. Air Force, Department of Defense, or U.S. Government.

Footnotes

1. ^The construction and training of RVFL networks is left as a “black box” procedure. How to best choose a specific activation function ρ(z) and train each RVFL network given by Equation 22 is outside of the scope of this study. The reader may, for instance, select from the range of methods available for training neural networks.

References

1. Krizhevsky A, Sutskever I, Hinton GE. Imagenet classification with deep convolutional neural networks. In: Advances in Neural Information Processing Systems. (2012). p. 1097–105.

Google Scholar

2. Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, et al. Going deeper with convolutions. In: Proceedings of the IEEE conference on computer vision and pattern recognition. (2015). p. 1–9. doi: 10.1109/CVPR.2015.7298594

Crossref Full Text | Google Scholar

3. He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision And Pattern Recognition. (2016). p. 770–778. doi: 10.1109/CVPR.2016.90

Crossref Full Text | Google Scholar

4. Huang G, Liu Z, Van Der Maaten L, Weinberger KQ. Densely connected convolutional networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. (2017). p. 4700–4708. doi: 10.1109/CVPR.2017.243

Crossref Full Text | Google Scholar

5. Yang Y, Zhong Z, Shen T, Lin Z. Convolutional neural networks with alternately updated clique. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. (2018). p. 2413–2422. doi: 10.1109/CVPR.2018.00256

Crossref Full Text | Google Scholar

6. Barron AR. Universal approximation bounds for superpositions of a sigmoidal function. IEEE Trans Inf Theory. (1993) 39:930–45. doi: 10.1109/18.256500

Crossref Full Text | Google Scholar

7. Candés EJ. Harmonic analysis of neural networks. Appl Comput Harm Analysis. (1999) 6:197–218. doi: 10.1006/acha.1998.0248

Crossref Full Text | Google Scholar

8. Vershynin R. Memory capacity of neural networks with threshold and ReLU activations. arXiv preprint arXiv:200106938. (2020).

Google Scholar

9. Baldi P, Vershynin R. The capacity of feedforward neural networks. Neural Netw. (2019) 116:288–311. doi: 10.1016/j.neunet.2019.04.009

PubMed Abstract | Crossref Full Text | Google Scholar

10. Huang GB, Babri HA. Upper bounds on the number of hidden neurons in feedforward networks with arbitrary bounded nonlinear activation functions. IEEE Trans Neural Netw. (1998) 9:224–9. doi: 10.1109/72.655045

PubMed Abstract | Crossref Full Text | Google Scholar

11. Suganthan PN. Letter: On non-iterative learning algorithms with closed-form solution. Appl Soft Comput. (2018) 70:1078–82. doi: 10.1016/j.asoc.2018.07.013

Crossref Full Text | Google Scholar

12. Olson M, Wyner AJ, Berk R. Modern Neural Networks Generalize on Small Data Sets. In: Proceedings of the 32Nd International Conference on Neural Information Processing Systems. NIPS'18. Curran Associates Inc. (2018). p. 3623–3632.

Google Scholar

13. Schmidt WF, Kraaijveld MA, Duin RPW. Feedforward neural networks with random weights. In: Proceedings 11th IAPR International Conference on Pattern Recognition. Vol. II. Conference B: Pattern Recognition Methodology and Systems. (1992). p. 1–4.

Google Scholar

14. Te Braake HAB, Van Straten G. Random activation weight neural net (RAWN) for fast non-iterative training. Eng Appl Artif Intell. (1995) 8:71–80. doi: 10.1016/0952-1976(94)00056-S

Crossref Full Text | Google Scholar

15. Pao YH, Takefuji Y. Functional-link net computing: theory, system architecture, and functionalities. Computer. (1992) 25:76–9. doi: 10.1109/2.144401

Crossref Full Text | Google Scholar

16. Igelnik B, Pao YH. Stochastic choice of basis functions in adaptive function approximation and the functional-link net. IEEE Trans Neur Netw. (1995) 6:1320–9. doi: 10.1109/72.471375

PubMed Abstract | Crossref Full Text | Google Scholar

17. Huang GB, Zhu QY, Siew CK. Extreme learning machine: theory and applications. Neurocomputing. (2006) 70:489–501. doi: 10.1016/j.neucom.2005.12.126

Crossref Full Text | Google Scholar

18. Pao YH, Park GH, Sobajic DJ. Learning and generalization characteristics of the random vector functional-link net. Neurocomputing. (1994) 6:163–80. doi: 10.1016/0925-2312(94)90053-1

Crossref Full Text | Google Scholar

19. Pao YH, Phillips SM. The functional link net and learning optimal control. Neurocomputing. (1995) 9:149–64. doi: 10.1016/0925-2312(95)00066-F

Crossref Full Text | Google Scholar

20. Chen CP, Wan JZ, A. rapid learning and dynamic stepwise updating algorithm for flat neural networks and the application to time-series prediction. IEEE Trans Syst Man Cybern B. (1999) 29:62–72. doi: 10.1109/3477.740166

PubMed Abstract | Crossref Full Text | Google Scholar

21. Park GH, Pao YH. Unconstrained word-based approach for off-line script recognition using density-based random-vector functional-link net. Neurocomputing. (2000) 31:45–65. doi: 10.1016/S0925-2312(99)00149-6

Crossref Full Text | Google Scholar

22. Zhang L, Suganthan PN. Visual tracking with convolutional random vector functional link network. IEEE Trans Cybern. (2017) 47:3243–53. doi: 10.1109/TCYB.2016.2588526

PubMed Abstract | Crossref Full Text | Google Scholar

23. Zhang L, Suganthan PN. Benchmarking ensemble classifiers with novel co-trained kernel ridge regression and random vector functional link ensembles [research frontier]. IEEE Comput Intell Mag. (2017) 12:61–72. doi: 10.1109/MCI.2017.2742867

Crossref Full Text | Google Scholar

24. Katuwal R, Suganthan PN, Zhang L. An ensemble of decision trees with random vector functional link networks for multi-class classification. Appl Soft Comput. (2018) 70:1146–53. doi: 10.1016/j.asoc.2017.09.020

Crossref Full Text | Google Scholar

25. Vuković N, Petrović M, Miljković Z. A comprehensive experimental evaluation of orthogonal polynomial expanded random vector functional link neural networks for regression. Appl Soft Comput. (2018) 70:1083–1096. doi: 10.1016/j.asoc.2017.10.010

Crossref Full Text | Google Scholar

26. Tang L, Wu Y, Yu L, A. non-iterative decomposition-ensemble learning paradigm using RVFL network for crude oil price forecasting. Appl Soft Comput. (2018) 70:1097–108. doi: 10.1016/j.asoc.2017.02.013

Crossref Full Text | Google Scholar

27. Dash Y, Mishra SK, Sahany S, Panigrahi BK. Indian summer monsoon rainfall prediction: A comparison of iterative and non-iterative approaches. Appl Soft Comput. (2018) 70:1122–34. doi: 10.1016/j.asoc.2017.08.055

Crossref Full Text | Google Scholar

28. Henríquez PA, Ruz G. Twitter sentiment classification based on deep random vector functional link. In: 2018 International Joint Conference on Neural Networks (IJCNN). IEEE (2018). p. 1–6. doi: 10.1109/IJCNN.2018.8489703

Crossref Full Text | Google Scholar

29. Katuwal R, Suganthan PN, Tanveer M. Random vector functional link neural network based ensemble deep learning. arXiv preprint arXiv:190700350. (2019).

PubMed Abstract | Google Scholar

30. Zhang Y, Wu J, Cai Z, Du B, Yu PS. An unsupervised parameter learning model for RVFL neural network. Neural Netw. (2019) 112:85–97. doi: 10.1016/j.neunet.2019.01.007

PubMed Abstract | Crossref Full Text | Google Scholar

31. Hornik K. Approximation capabilities of multilayer feedforward networks. Neural Netw. (1991) 4:251–7. doi: 10.1016/0893-6080(91)90009-T

Crossref Full Text | Google Scholar

32. Leshno M, Lin VY, Pinkus A, Schocken S. Multilayer feedforward networks with a nonpolynomial activation function can approximate any function. Neural Netw. (1993) 6:861–7. doi: 10.1016/S0893-6080(05)80131-5

Crossref Full Text | Google Scholar

33. Li JY, Chow WS, Igelnik B, Pao YH. Comments on “Stochastic choice of basis functions in adaptive function approximation and the functional-link net” [with reply]. IEEE Trans Neur Netw. (1997) 8:452–4. doi: 10.1109/72.557702

Crossref Full Text | Google Scholar

34. Burkhardt DB, San Juan BP, Lock JG, Krishnaswamy S, Chaffer CL. Mapping phenotypic plasticity upon the cancer cell state landscape using manifold learning. Cancer Discov. (2022) 12:1847–59. doi: 10.1158/2159-8290.CD-21-0282

PubMed Abstract | Crossref Full Text | Google Scholar

35. Mitchell-Heggs R, Prado S, Gava GP, Go MA, Schultz SR. Neural manifold analysis of brain circuit dynamics in health and disease. J Comput Neurosci. (2023) 51:1–21. doi: 10.1007/s10827-022-00839-3

PubMed Abstract | Crossref Full Text | Google Scholar

36. Dick J, Kuo FY, Sloan IH. High-dimensional integration: the quasi-Monte Carlo way. Acta Numerica. (2013) 22:133–288. doi: 10.1017/S0962492913000044

Crossref Full Text | Google Scholar

37. Ledoux M. The Concentration of Measure Phenomenon. Providence, RI: American Mathematical Soc. (2001).

Google Scholar

38. Massart P. About the constants in Talagrand's deviation inequalities for empirical processes. Technical Report, Laboratoire de statistiques, Universite Paris Sud. (1998).

Google Scholar

39. Talagrand M. New concentration inequalities in product spaces. Invent Mathem. (1996) 126:505–63. doi: 10.1007/s002220050108

Crossref Full Text | Google Scholar

40. Shaham U, Cloninger A, Coifman RR. Provable approximation properties for deep neural networks. Appl Comput Harmon Anal. (2018) 44:537–57. doi: 10.1016/j.acha.2016.04.003

Crossref Full Text | Google Scholar

41. Tu LW. An Introduction to Manifolds. New York: Springer. (2010). doi: 10.1007/978-1-4419-7400-6_3

Crossref Full Text | Google Scholar

42. Allard WK, Chen G, Maggioni M. Multi-scale geometric methods for data sets II: geometric multi-resolution analysis. Appl Comput Harmon Anal. (2012) 32:435–62. doi: 10.1016/j.acha.2011.08.001

Crossref Full Text | Google Scholar

43. Iwen MA, Krahmer F, Krause-Solberg S, Maly J. On recovery guarantees for one-bit compressed sensing on manifolds. Discr Comput Geom. (2018) 65:953–998. doi: 10.1109/SAMPTA.2017.8024465

Crossref Full Text | Google Scholar

44. Maggioni M, Minsker S, Strawn N. Multiscale dictionary learning: non-asymptotic bounds and robustness. J Mach Learn Res. (2016) 17:43–93.

Google Scholar

45. Liao W, Maggioni M. Adaptive Geometric Multiscale Approximations for Intrinsically Low-dimensional Data. J Mach Learn Res. (2019) 20:1–63.

Google Scholar

46. Stein EM, Weiss G. Introduction to Fourier Analysis on Euclidean Spaces. London: Princeton University Press (1971). doi: 10.1515/9781400883899

Crossref Full Text | Google Scholar

47. Rudin W. Functional Analysis. New York: McGraw-Hill (1991).

Google Scholar

Keywords: machine learning, feed-forward neural networks, function approximation, smooth manifold, random vector functional link

Citation: Needell D, Nelson AA, Saab R, Salanevich P and Schavemaker O (2024) Random vector functional link networks for function approximation on manifolds. Front. Appl. Math. Stat. 10:1284706. doi: 10.3389/fams.2024.1284706

Received: 28 August 2023; Accepted: 21 March 2024;
Published: 17 April 2024.

Edited by:

HanQin Cai, University of Central Florida, United States

Reviewed by:

Jingyang Li, The University of Hong Kong, Hong Kong SAR, China
Dong Xia, Hong Kong University of Science and Technology, Hong Kong SAR, China

Copyright © 2024 Needell, Nelson, Saab, Salanevich and Schavemaker. This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.

*Correspondence: Deanna Needell, ZGVhbm5hQG1hdGgudWNsYS5lZHU=

Disclaimer: All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Random vector functional link networks for function approximation on manifolds

1 Introduction

1.1 Notation

1.2 Main results

1.3 Organization

2 Materials and methods

2.1 A concentration bound for classic Monte-Carlo integration

2.2 Smooth, compact manifolds in Euclidean space

3 Results

3.1 Proof of Theorem 1

3.1.1 Proof of Theorem 1 when ρ ∈ L1(ℝ)∩L∞(ℝ)

3.1.2 Proof of Theorem 1 when ρ′ ∈ L1(ℝ)∩L∞(ℝ)

3.2 Proof of Theorem 2

3.3 Results on sub-manifolds of Euclidean space

3.3.1 Adapting RVFL networks to d-manifolds

3.3.2 Main results on d-manifolds

3.4 Numerical simulations

3.5 Proofs of technical lemmas

3.5.1 Proof of Lemma 3

3.5.2 Proof of Lemma 4: the limit-integral representation

3.5.3 Proof of Lemma 5: Monte-Carlo integral approximation

3.5.4 Proof of Theorem 1 when ρ′ ∈ L1(ℝ)∩L∞(ℝ)

3.5.5 Proof of Theorem 8

4 Discussion

Data availability statement

Author contributions

Funding

Acknowledgments

Conflict of interest

Publisher's note

Author disclaimer

Footnotes

References

3.1.1 Proof of Theorem 1 when ρ ∈ L¹(ℝ)∩L^∞(ℝ)

3.1.2 Proof of Theorem 1 when ρ′ ∈ L¹(ℝ)∩L^∞(ℝ)

3.5.4 Proof of Theorem 1 when ρ′ ∈ L¹(ℝ)∩L^∞(ℝ)