GoGuides Verified Text
DIFFERENTIAL EQUATION
SHA-256 integrity check: match
Source
Encyclopaedia Britannica (1911) / britannica_1911
License
public_domain
Chunk ID
1911:differential equation:5e979facfada
Section
Hash Algorithm
sha256
Stored Hash
b6400e9f7a0abdcbd69a3c48c4c3d7edc2c96eed107dffcd211a737b1d2cabf2
Computed Hash
b6400e9f7a0abdcbd69a3c48c4c3d7edc2c96eed107dffcd211a737b1d2cabf2
Normalizer
ggnorm 1.0
Observed
2026-02-08 18:42:30
Source URL
Verified Text
differential equation, in mathematics, a relation between one or more functions and their differential coefficients. the subject is treated here in two parts: (1) an elementary introduction dealing with the more commonly recognized types of differential equations which can be solved by rule; and (2) the general theory. _part i.--elementary introduction._ of equations involving only one independent variable, x (known as _ordinary_ differential equations), and one dependent variable, y, and containing only the first differential coefficient dy/dx (and therefore said to be of the first _order_), the simplest form is that reducible to the type dy/dx = f(x)/f(y), leading to the result ff(y)dy - ff(x)dx = a, where a is an arbitrary constant; this result is said to solve the differential equation, the problem of evaluating the integrals belonging to the integral calculus. another simple form is dy/dx + yp = q, where p, q are functions of x only; this is known as the linear equation, since it contains y and dy/dx only to the first degree. if fpdx = u, we clearly have d /dy \ --(ye^u) =e^u ( -- + py) = e^u q, dx \dx / so that y = e^-u(fe^u qdx + a) solves the equation, and is the only possible solution, a being an arbitrary constant. the rule for the solution of the linear equation is thus to multiply the equation by e^u, where u = fpdx. a third simple and important form is that denoted by y = px + f(p), where p is an abbreviation for dy/dx; this is known as clairaut's form. by differentiation in regard to x it gives dp dp p = p + x-- + f'(p)--, dx dx where d f'(p) = -- f(p); dp thus, either (i.) dp/dx = 0, that is, p is constant on the curve satisfying the differential equation, which curve is thus any one of the straight lines y = cx = f(c), where c is an arbitrary constant, or else, (ii.) x + [f]'(p) = 0; if this latter hypothesis be taken, and p be eliminated between x + f'(p) = 0 and y = px + f(p), a relation connecting x and y, not containing an arbitrary constant, will be found, which obviously represents the envelope of the straight lines y = cx + f(c). in general if a differential equation [phi](x, y, dy/dx) = 0 be satisfied by any one of the curves f(x, y, c) = 0, where c is an arbitrary constant, it is clear that the envelope of these curves, when existent, must also satisfy the differential equation; for this equation prescribes a relation connecting only the co-ordinates x, y and the differential coefficient dy/dx, and these three quantities are the same at any point of the envelope for the envelope and for the particular curve of the family which there touches the envelope. the relation expressing the equation of the envelope is called a _singular_ solution of the differential equation, meaning an _isolated_ solution, as not being one of a family of curves depending upon an arbitrary parameter. an extended form of clairaut's equation expressed by y = xf(p) + f(p) may be similarly solved by first differentiating in regard to p, when it reduces to a linear equation of which x is the dependent and p the independent variable; from the integral of this linear equation, and the original differential equation, the quantity p is then to be eliminated. other types of solvable differential equations of the first order are (1) m dy/dx = n, where m, n are homogeneous polynomials in x and y, of the same order; by putting v = y/x and eliminating y, the equation becomes of the first type considered above, in v and x. an equation (ab <> ba) (ax + by + c)dy/dx = ax + by + c may be reduced to this rule by first putting x + h, y + k for x and y, and determining h, k so that ah + bk + c = 0, ah + bk + c = 0. (2) an equation in which y does not explicitly occur, f(x, dy/dx) = 0, may, theoretically, be reduced to the type dy/dx = f(x); similarly an equation f(y, dy/dx) = 0. (3) an equation f(dy/dx, x, y) = 0, which is an integral polynomial in dy/dx, may, theoretically, be solved for dy/dx, as an algebraic equation; to any root dy/dx = f1(x, y) corresponds, suppose, a solution [phi]1(x, y, c) = 0, where c is an arbitrary constant; the product equation [phi]1(x, y, c)[phi]2(x, y, c) ... = 0, consisting of as many factors as there were values of dy/dx, is effectively as general as if we wrote [phi]1(x, y, c1) [phi]2(x, y, c2) ... = 0; for, to evaluate the first form, we must necessarily consider the factors separately, and nothing is then gained by the multiple notation for the various arbitrary constants. the equation [phi]1(x, y, c)[phi]2(x, y, c) ... = 0 is thus the solution of the given differential equation. in all these cases there is, except for cases of singular solutions, one and only one arbitrary constant in the most general solution of the differential equation; that this must necessarily be so we may take as obvious, the differential equation being supposed to arise by elimination of this constant from the equation expressing its solution and the equation obtainable from this by differentiation in regard to x. a further type of differential equation of the first order, of the form dy/dx = a + by + cy2 in which a, b, c are functions of x, will be briefly considered below under differential equations of the second order. when we pass to ordinary differential equations of the second order, that is, those expressing a relation between x, y, dy/dx and d2y/dx2, the number of types for which the solution can be found by a known procedure is very considerably reduced. consider the general linear equation d2y dy --- + p-- + qy = r, dx2 dx where p, q, r are functions of x only. there is no method always effective; the main general result for such a linear equation is that if any particular function of x, say y1, can be discovered, for which d2y1 dy1 ---- + p--- + qy1 = 0, dx2 dx then the substitution y = y1[eta] in the original equation, with r on the right side, reduces this to a linear equation of the first order with the dependent variable d[eta]/dx. in fact, if y = y1[eta] we have dy d[eta] dy1 d2y d2[eta] dy1 d[eta] d2y1 -- = y1------ + [eta]--- and --- = y1------- + 2--- ------ + [eta]----, dx dx dx dx2 dx2 dx dx dx2 and thus d2y dy d2[eta] / dy1 \ d[eta] /d2y1 dy1 \ --- + p -- + qy = y1------- + ( 2--- + py1) ------ + ( ---- + p--- + qy1)[eta]; dx2 dx dx2 \ dx / dx \ dx2 dx / if then d2y1 dy1 ---- + p --- + qy1 = 0, dx2 dx and z denote d[eta]/dx, the original differential equation becomes dz / dy1 \ y1-- + ( 2--- + py1)z = r. dx \ dx / from this equation z can be found by the rule given above for the linear equation of the first order, and will involve one arbitrary constant; thence y = y1 [eta] = y1 [int] zdx + ay1, where a is another arbitrary constant, will be the general solution of the original equation, and, as was to be expected, involves two arbitrary constants. the case of most frequent occurrence is that in which the coefficients p, q are constants; we consider this case in some detail. if [t]* be a root of the quadratic equation [t]2 + [t]p + q = 0, it can be at once seen that a particular integral of the differential equation with zero on the right side is y1 = e^[theta]x. supposing first the roots of the quadratic equation to be different, and [phi] to be the other root, so that [p] + [t] = -p, the auxiliary differential equation for z, referred to above, becomes dz/dx + ([t] - [p])z = re^(-[t]^x), which leads to ze^{([t]-[p])^x} = b + [int] re^(-[p]^x)dx, where b is an arbitrary constant, and hence to (*) [t] = [theta]; [p] = [phi]. _ _ _ / / / y = ae^([t]^x) + e^([t]^x)| be^([p]-[t])^x dx + e^[t]^x | e^([p]-[t])^x | re^-[p]^x dxdx, _/ _/ _/ or say to y = ae^[t]^x + ce^[p]^x + u, where a, c are arbitrary constants and u is a function of x, not present at all when r = 0. if the quadratic equation [t]2 + p[t] + q = 0 has equal roots, so that 2[t] = -p, the auxiliary equation in z becomes dz/dx = re^-[t]^x, giving z = b + [int] re^-[t]^x dx, where b is an arbitrary constant, and hence _ _ / / y = (a + bx)e^[t]^x + e^[t]^x | | re^-[t]^x dxdx, _/ _/ or, say, y = (a + bx)e^[t]^x + u, where a, b are arbitrary constants, and u is a function of x not present at all when r = 0. the portion ae^[t]^x + be^[p]^x or (a + bx)e^[t]^x of the solution, which is known as the _complementary function_, can clearly be written down at once by inspection of the given differential equation. the remaining portion u may, by taking the constants in the complementary function properly, be replaced by any particular solution whatever of the differential equation d2v dy --- + p -- + qy = r; dx2 dx for if u be any particular solution, this has a form u = a0 e^[t]^x + b0 e^[p]^x + u, or a form u = (a0 + b0x)e^[t]^x + u; thus the general solution can be written (a - a0)e^[t]^x + (b - b0)e^[p]^x + u, or {a - a0 + (b - b0)x}e^[t]^x + u, where a - a0, b - b0, like a, b, are arbitrary constants. a similar result holds for a linear differential equation of any order, say d^n y d^n-1 y ----- + p1 ------- + ... + p_n y = r, dx_n dx^n-1 where p1, p2, ... pn are constants, and r is a function of x. if we form the algebraic equation [t]^n + p1[t]^n-1 + ... + p_n = 0, and all the roots of this equation be different, say they are [t]1, [t]2, ... [t]n, the general solution of the differential equation is y = a1 e^[t]1^x + a2 e^[t]2^x + ... + a_n e^[t]_n^x + u, where a1, a2, ... an are arbitrary constants, and u is any particular solution whatever; but if there be one root [t]1 repeated r times, the terms a1 e^[t]1^x + ... + a_r e^[t]_r^x must be replaced by (a1 + a2x + ... + a_r x^r-1)e^[t]1x where a1, ... an are arbitrary constants; the remaining terms in the complementary function will similarly need alteration of form if there be other repeated roots. to complete the solution of the differential equation we need some method of determining a particular integral u; we explain a procedure which is effective for this purpose in the cases in which r is a sum of terms of the form e^ax[p](x), where [p](x) is an integral polynomial in x; this includes cases in which r contains terms of the form cos bx·[p](x) or sin bx·[p](x). denote d/dx by d; it is clear that if u be any function of x, d(e^ax u) = e^ax du + ae^ax u, or say, d(e^ax u) = e^ax (d + a)u; hence d2(e^ax u), i.e. d2/dx2 (e^ax u), being equal to d(e^ax v), where v=(d + a)u, is equal to e^ax(d + a)v, that is to e^ax(d + a)2u. in this way we find d^n(e^ax u) = e^ax(d + a)^n u, where n is any positive integer. hence if [psi](d) be any polynomial in d with constant coefficients, [psi](d)(e^ax u) = e^ax [psi](d + a)u. next, denoting [int] udx by d^-1 u, and any solution of the differential equation dz/dx + az = u by z = (d + a)^-1 u, we have d[e^ax(d + a)^-1 u] = d(e^ax z) = e^ax(d + a)z = e^ax u, so that we may write d^-1(e^ax u) = e^ax(d+a)^-1 u, where the meaning is that one value of the left side is equal to one value of the right side; from this, the expression d^-2(e^axu), which means d^-1[d^-1(e^ax u)], is equal to d^-1(e^ax z) and hence to e^ax(d + a)^-1 z, which we write e^ax(d + a)^-2 u; proceeding thus we obtain d^-n(e^ax u) = e^ax(d + a)^-n u, where n is any positive integer, and the meaning, as before, is that one value of the first expression is equal to one value of the second. more generally, if [psi](d) be any polynomial in d with constant coefficients, and we agree to denote by 1/[psi](d) u any solution z of the differential equation [psi](d)z = u, we have, if v = 1/[psi](d + a) u, the identity [psi](d)(e^ax v) = e^ax [psi](d + a)v = e^ax u, which we write in the form 1 1 --------(e^ax u) = e^ax ------------ u. [psi](d) [psi](d + a) this gives us the first step in the method we are explaining, namely that a solution of the differential equation [psi](d)y = e^ax u + e^bx v + ... where u, v, ... are any functions of x, is any function denoted by the expression 1 1 e^ax ------------ u + e^ax ------------ v + .... [psi](d + a) [psi](d + b) it is now to be shown how to obtain one value of 1/[psi](d + a) u, when u is a polynomial in x, namely one solution of the differential equation [psi](d + a)z = u. let the highest power of x entering in u be x^m; if t were a variable quantity, the rational fraction in t, 1/[psi](t + a), by first writing it as a sum of partial fractions, or otherwise, could be identically written in the form k_r t^-r + k_r-1 t^-r+1 + ... + k1 t^-1 + h + h1t + ... + h_m t^m + t^m+1 [p](t)/[psi](t + a), where [p](t) is a polynomial in t; this shows that there exists an identity of the form 1 = [psi](t + a)(k_r t^-r + ... + k1t^-1 + h + h1t + ... + h_m t^m) + [p](t)t^m+1, and hence an identity u = [psi](d + a)[k_r d^-r + ... + k1d^-1 + h + h1d + ... + h_m d^m]u + [p](d)d^m+1 u; in this, since u contains no power of x higher than x^m, the second term on the right may be omitted. we thus reach the conclusion that a solution of the differential equation [psi](d + a)z = u is given by z = (k_r d^-r + ... + k1d^-1 + h + h1d + ... + h_m d^m)u, of which the operator on the right is obtained simply by expanding 1/[psi](d + a) in ascending powers of d, as if d were a numerical quantity, the expansion being carried as far as the highest power of d which, operating upon u, does not give zero. in this form every term in z is capable of immediate calculation. _example._--for the equation d^4v d2y ---- + 2--- + y = x3 cos x or (d2 + 1)2y = x3 cos x, dx^4 dx3 the roots of the associated algebraic equation ([t]2+1)2 = 0 are [t] = ±i, each repeated; the complementary function is thus (a + bx)e^ix + (c + dx)e^ix, where a, b, c, d are arbitrary constants; this is the same as (h + kx) cos x + (m + nx) sin x, where h, k, m, n are arbitrary constants. to obtain a particular integral we must find a value of (1 + d2)^-2 x3 cos x; this is the real part of (1+d2)^-2 e^ix x3 and hence of e^ix [1 + (d + i)2]^-2 x3 or e^ix [2id(1 + 1⁄2id)]^-2 x3, or -1⁄4e^ix d^-2 (1 + id - 3⁄4d2 - 1⁄2id3 + 5/16 d^4 + 3/16 id^5 ...)x3, or -1⁄4e^ix(1/20 x^5 + 1⁄4ix^4 - 3⁄4x3 - 3/2 ix2 + 15/8 x + 9/8 i); the real part of this is -1⁄4(1/20 x^5 - 3⁄4x2 + 15/8 x) cos x + 1⁄4(1⁄4x^4 - 3/2 x2 + 9/8) sin x. this expression added to the complementary function found above gives the complete integral; and no generality is lost by omitting from the particular integral the terms -15/32 x cos x + 9/32 sin x, which are of the types of terms already occurring in the complementary function. the symbolical method which has been explained has wider applications than that to which we have, for simplicity of explanation, restricted it. for example, if [psi](x) be any function of x, and a1, a2, ... an be different constants, and [(t + a1) (t + a2) ... (t + an)]^-1 when expressed in partial fractions be written [sigma]c_m(t + a_m)^-1, a particular integral of the differential equation (d + a1)(d + a2) ... (d + a_n)y = [psi](x) is given by y = [sigma]c_m(d + a_m)^-1 [psi](x) = [sigma]c_m(d + a_m)^-1 e^-a m^x e^a m^x [psi](x) = [sigma]c_m e^-a m^x d^-1 (e^a m^x [psi](x)) = [sigma]c_m e^-a m^x [int] e^a m^x [psi](x)dx. the particular integral is thus expressed as a sum of n integrals. a linear differential equation of which the left side has the form d^ny d^n-1 y dy x^n ---- + p1x^n-1 ------- + ... + p_n-1 x-- + p_n y, dx^n dx^n-1 dx where p1, ... pn are constants, can be reduced to the case considered above. writing x = e^t we have the identity d^mu x^m ---- = [t]([t] - 1)([t] - 2) ... ([t] - m + 1)u, where [t] = d/dt. dx^m when the linear differential equation, which we take to be of the second order, has variable coefficients, though there is no general rule for obtaining a solution in finite terms, there are some results which it is of advantage to have in mind. we have seen that if one solution of the equation obtained by putting the right side zero, say y1, be known, the equation can be solved. if y2 be another solution of d2y dy --- + p-- + qy = 0, dx2 dx there being no relation of the form my1 + ny2 = k, where m, n, k are constants, it is easy to see that d/dx(y1'y2 - y1y2') = p(y1'y2 - y1y2'), so that we have y1'y2 - y1y2' = a exp.([int] pdx), where a is a suitably chosen constant, and exp. z denotes e^z. in terms of the two solutions y1, y2 of the differential equation having zero on the right side, the general solution of the equation with r = [phi](x) on the right side can at once be verified to be ay1 + by2 + y1u - y2v, where u, v respectively denote the integrals _ _ / / u = |y2[phi](x)(y1'y2 - y2'y1)^-1 dx, v = |y1[phi](x)(y1'y2 - y2'y1)^-1 dx. _/ _/ the equation d2y dy --- + p-- + qy = 0, dx2 dx by writing y = v exp. (-1⁄2 [int] pdx), is at once seen to be reduced to d2v/dx2 + 1v = 0, where 1 = q - 1⁄2dp/dx - 1⁄4p2. if [eta] = - 1/v dv/dx, the equation d2v/dx2 + 1v = 0 becomes d[eta]/dx = 1 + [eta]2, a non-linear equation of the first order. more generally the equation d[eta] ------ = a + b[eta] + c[eta]2, dx where a, b, c are functions of x, is, by the substitution 1 dy [eta] = - -- --, cy dx reduced to the linear equation d2y / 1 dc\ dy --- - ( b + - -- )-- + acy = 0. dx2 \ c dx/ dx the equation d[eta] ------ = a + b[eta] + c[eta]2, dx known as riccati's equation, is transformed into an equation of the same form by a substitution of the form [eta] = (ay + b)/(cy + d), where a, b, c, d are any functions of x, and this fact may be utilized to obtain a solution when a, b, c have special forms; in particular if any particular solution of the equation be known, say [eta]0, the substitution [eta] = [eta]0 - 1/y enables us at once to obtain the general solution; for instance, when d /a\ 2b = -- log( - ), dx \c/ a particular solution is [eta]0 = [root](-a/c). this is a case of the remark, often useful in practice, that the linear equation d2y d[phi] dy [phi](x)--- + 1⁄2------ -- + [mu]y = 0, dx2 dx dx where [mu] is a constant, is reducible to a standard form by taking a new independent variable _ / z = | dx[[p](x)]^-1⁄2. _/ we pass to other types of equations of which the solution can be obtained by rule. we may have cases in which there are two dependent variables, x and y, and one independent variable t, the differential coefficients dx/dt, dy/dt being given as functions of x, y and t. of such equations a simple case is expressed by the pair dx dy -- = ax + by + c, -- = a'x + b'y + c', dt dt wherein the coefficients a, b, c, a', b', c', are constants. to integrate these, form with the constant [lambda] the differential coefficient of z = x + [lambda]y, that is dz/dt = (a + [lambda]a')x + (b + [lambda]b')y + c + [lambda]c', the quantity [lambda] being so chosen that b + [lambda]b' = [lambda](a + [lambda]a'), so that we have dz/dt = (a + [lambda]a')z + c + [lambda]c'; this last equation is at once integrable in the form z(a + [lambda]a') + c + [lambda]c' = ae^(a + [lambda]a')t, where a is an arbitrary constant. in general, the condition b + [lambda]b' = [lambda](a + [lambda]a') is satisfied by two different values of [lambda], say [lambda]1, [lambda]2; the solutions corresponding to these give the values of x +[lambda]1y and x + [lambda]2y, from which x and y can be found as functions of t, involving two arbitrary constants. if, however, the two roots of the quadratic equation for [lambda] are equal, that is, if (a - b')2 + 4a'b = 0, the method described gives only one equation, expressing x + [lambda]y in terms of t; by means of this equation y can be eliminated from dx/dt = ax + by + c, leading to an equation of the form dx/dt = px + q + re^(a + [lambda]a')t, where p, q, r are constants. the integration of this gives x, and thence y can be found. a similar process is applicable when we have three or more dependent variables whose differential coefficients in regard to the single independent variables are given as linear functions of the dependent variables with constant coefficients. another method of solution of the equations dx/dt = ax + by + c, dy/dt = a'x + b'y + c', consists in differentiating the first equation, thereby obtaining d2x dx dy --- = a-- + b--; dt2 dt dx from the two given equations, by elimination of y, we can express dy/dt as a linear function of x and dx/dt; we can thus form an equation of the shape d2x/dt2 = p + qx + rdx/dt, where p, q, r are constants; this can be integrated by methods previously explained, and the integral, involving two arbitrary constants, gives, by the equation dx/dt = ax + by + c, the corresponding value of y. conversely it should be noticed that any single linear differential equation d2x dx --- = u + vx + w--, dt2 dt where u, v, w are functions of t, by writing y for dx/dt, is equivalent with the two equations dx/dt = y, dy/dt = u + vx + wy. in fact a similar reduction is possible for any system of differential equations with one independent variable. equations occur to be integrated of the form xdx + ydy + zdz = 0, where x, y, z are functions of x, y, z. we consider only the case in which there exists an equation [phi](x, y, z) = c whose differential dp[phi] dp[phi] dp[phi] -------dx + -------dy + -------dz = 0 dpx dpy dpz is equivalent with the given differential equation; that is, [mu] being a proper function of x, y, z, we assume that there exist equations dp[phi] dp[phi] v[phi] ------- = [mu]x, ------- = [mu]y, ------ = [mu]z; dpx vy vz these equations require dp dp ---([mu]y) = ---([mu]z), &c., dpz dpy and hence /dpz dpy\ /dpx dpz\ /dpy dpx\ x( --- - --- ) + y( --- - --- ) + z( --- - --- ) = 0; \dpy dpz/ \dpz dpx/ \dpx dpy/ conversely it can be proved that this is sufficient in order that [mu] may exist to render [mu](xdx + ydy + zdz) a perfect differential; in particular it may be satisfied in virtue of the three equations such as dpz dpy --- - --- = 0; dpy dpz in which case we may take [mu] = 1. assuming the condition in its general form, take in the given differential equation a plane section of the surface [phi] = c parallel to the plane z, viz. put z constant, and consider the resulting differential equation in the two variables x, y, namely xdx + ydy = 0; let [psi](x, y, z) = constant, be its integral, the constant z entering, as a rule, in [psi] because it enters in x and y. now differentiate the relation [psi](x, y, z) = [f](z), where [f] is a function to be determined, so obtaining dp[psi] dp[psi] /dp[psi] df\ -------dx + -------dy + ( ------- - -- )dz = 0; dpx dpy \ dpz dz/ there exists a function [sigma] of x, y, z such that dp[psi] dp[psi] -------- = [sigma]x, ------- = [sigma]y, dpx dpy because [psi] = constant, is the integral of xdx + ydy = 0; we desire to prove that [f] can be chosen so that also, in virtue of [psi](x, y, z) = f(z), we have dp[psi] df df dp[psi] ------- - -- = [sigma]z, namely -- = ------- - [sigma]z; dpz dz dz dpz if this can be proved the relation [psi](x, y, z) - f(z) = constant, will be the integral of the given differential equation. to prove this it is enough to show that, in virtue of [psi](x, y, z) = [f](z), the function dp[psi]/dpx - [sigma]z can be expressed in terms of z only. now in consequence of the originally assumed relations, dp[psi] dp[phi] dp[phi] ------- = [mu]x, ------- = [mu]y, ------- = [mu]z, dpx dpy dpz we have dp[psi] /dp[phi] [sigma] dp[psi] /dp[phi] ------- / ------- = ------- = ------- / -------, dpx / dpx [mu] dpy / dpy and hence dp[psi] dp[phi] dp[psi] dp[phi] ------- ------- - ------- ------- = 0; dpx dpy dpy dpx this shows that, as functions of x and y, [psi] is a function of [phi] (see the note at the end of part i. of this article, on jacobian determinants), so that we may write [psi] = f(z, [phi]), from which [sigma] dpf dp[psi] dpf dpf dp[phi] dpf [sigma] dpf ------- = -------; then ------- = --- + ------- ------- = --- + ------- · [mu]z = --- + [sigma]z [mu] dp[phi] dpz dpz dp[phi] dpz dpz [mu] dpz dp[psi] dpf or ------- - [sigma]z = ---; dpz dpz in virtue of [psi](x, y, z) = f(z), and [psi] = f(z, [phi]), the function [phi] can be written in terms of z only, thus dpf/dpz can be written in terms of z only, and what we required to prove is proved. consider lastly a simple type of differential equation containing _two_ independent variables, say x and y, and one dependent variable z, namely the equation dpz dpz p--- + q--- = r, dpx dpy where p, q, r are functions of x, y, z. this is known as lagrange's linear partial differential equation of the first order. to integrate this, consider first the ordinary differential equations dx/dz = p/r, dy/dz = q/r, and suppose that two functions u, v, of x, y, z can be determined, independent of one another, such that the equations u = a, v = b, where a, b are arbitrary constants, lead to these ordinary differential equations, namely such that dpu dpu dpu dpv dpv dpv p--- + q--- = r--- = 0 and p--- + q--- = r--- = 0. dpx dpy dpz dpx dpy dpz then if f(x, y, z) = 0 be a relation satisfying the original differential equations, this relation giving rise to dpf dpf dpz dpf dpf dpz dpf dpf dpf --- + --- --- = 0 and --- + --- --- = 0, we have p--- + q--- = r--- = 0. dpx dpz dpx dpy dpz dpy dpx dpy dpz it follows that the determinant of three rows and columns vanishes whose first row consists of the three quantities dpf/dpx, dpf/dpy, dpf/dpz, whose second row consists of the three quantities dpu/dpx, dpu/dpy, dpu/dpz, whose third row consists similarly of the partial derivatives of v. the vanishing of this so-called jacobian determinant is known to imply that f is expressible as a function of u and v, unless these are themselves functionally related, which is contrary to hypothesis (see the note below on jacobian determinants). conversely, any relation [phi](u, v) = 0 can easily be proved, in virtue of the equations satisfied by u and v, to lead to dz dz p-- + q-- = r. dx dx the solution of this partial equation is thus reduced to the solution of the two ordinary differential equations expressed by dx/p = dy/q = dz/r. in regard to this problem one remark may be made which is often of use in practice: when one equation u = a has been found to satisfy the differential equations, we may utilize this to obtain the second equation v = b; for instance, we may, by means of u = a, eliminate z--when then from the resulting equations in x and y a relation v = b has been found containing x and y and a, the substitution a = u will give a relation involving x, y, z. _note on jacobian determinants._--the fact assumed above that the vanishing of the jacobian determinant whose elements are the partial derivatives of three functions f, u, v, of three variables x, y, z, involves that there exists a functional relation connecting the three functions f, u, v, may be proved somewhat roughly as follows:-- the corresponding theorem is true for any number of variables. consider first the case of two functions p, q, of two variables x, y. the function p, not being constant, must contain one of the variables, say x; we can then suppose x expressed in terms of y and the function p; thus the function q can be expressed in terms of y and the function p, say q = q(p, y). this is clear enough in the simplest cases which arise, when the functions are rational. hence we have dpq dpq dpp dpq dpq dpp dpq --- = --- --- and --- = --- --- + ---; dpx dpp dpx dpy dpp dpy dpy these give dpp dpq dpp dpq dpp dpq --- --- - --- --- = --- ---; dpx dpy dpy dpx dpx dpy by hypothesis dpp/dpx is not identically zero; therefore if the jacobian determinant of p and q in regard to x and y is zero identically, so is dpq/dpy, or q does not contain y, so that q is expressible as a function of p only. conversely, such an expression can be seen at once to make the jacobian of p and q vanish identically. passing now to the case of three variables, suppose that the jacobian determinant of the three functions f, u, v in regard to x, y, z is identically zero. we prove that if u, v are not themselves functionally connected, f is expressible as a function of u and v. suppose first that the minors of the elements of dpf/dpx, dpf/dpy, dpf/dpz in the determinant are all identically zero, namely the three determinants such as dpu dpv dpu dpv --- --- - --- ---; dpy dpz dpz dpy then by the case of two variables considered above there exist three functional relations. [psi]1(u, v, x) = 0, [psi]2(u, v, y) = 0, [psi]3(u, v, z) = 0, of which the first, for example, follows from the vanishing of dpu dpv dpu dpv --- --- - --- ---. dpy dpz dpz dpy we cannot assume that x is absent from [psi]1, or y from [psi]2, or z from [psi]3; but conversely we cannot simultaneously have x entering in [psi]1, and y in [psi]2, and z in [psi]3, or else by elimination of u and v from the three equations [psi]1 = 0, [psi]2 = 0, [psi]3 = 0, we should find a necessary relation connecting the three independent quantities x, y, z; which is absurd. thus when the three minors of dpf/dpx, dpf/dpy, dpf/dpz in the jacobian determinant are all zero, there exists a functional relation connecting u and v only. suppose no such relation to exist; we can then suppose, for example, that dpu dpv dpu dpv --- --- - --- --- dpy dpz dpz dpy is not zero. then from the equations u(x, y, z) = u, v(x, y, z) = v we can express y and z in terms of u, v, and x (the attempt to do this could only fail by leading to a relation connecting u, v and x, and the existence of such a relation would involve that the determinant dpu dpv dpu dpv --- --- - --- --- dpy dpz dpz dpy was zero), and so write f in the form f(x, y, z) = [phi](u, v, x). we then have dpf dp[phi] dpu dp[phi] dpv dp[phi] dpf dp[phi] dpu dp[phi] dpv dpf dp[phi] dpu dp[phi] dpv --- = ------- --- + ------- --- + -------, --- = ------- --- + ------- ---, --- = ------- --- + ------- ---; dpx dpu dpx dpv dpx dpx dpy dpu dpy dpv dpy dpz dpu dpz dpv dpz thereby the jacobian determinant of f, u, v is reduced to dp[phi] /dpu dpv dpu dpv\ -------( --- --- - --- --- ); dpx \dpy dpz dpz dpy/ by hypothesis the second factor of this does not vanish identically; hence dp[phi]/dpx = 0 identically, and [phi] does not contain x; so that f is expressible in terms of u, v only; as was to be proved. _part ii.--general theory._ differential equations arise in the expression of the relations between quantities by the elimination of details, either unknown or regarded as unessential to the formulation of the relations in question. they give rise, therefore, to the two closely connected problems of determining what arrangement of details is consistent with them, and of developing, apart from these details, the general properties expressed by them. very roughly, two methods of study can be distinguished, with the names transformation-theories, function-theories; the former is concerned with the reduction of the algebraical relations to the fewest and simplest forms, eventually with the hope of obtaining explicit expressions of the dependent variables in terms of the independent variables; the latter is concerned with the determination of the general descriptive relations among the quantities which are involved by the differential equations, with as little use of algebraical calculations as may be possible. under the former heading we may, with the assumption of a few theorems belonging to the latter, arrange the theory of partial differential equations and pfaff's problem, with their geometrical interpretations, as at present developed, and the applications of lie's theory of transformation-groups to partial and to ordinary equations; under the latter, the study of linear differential equations in the manner initiated by riemann, the applications of discontinuous groups, the theory of the singularities of integrals, and the study of potential equations with existence-theorems arising therefrom. in order to be clear we shall enter into some detail in regard to partial differential equations of the first order, both those which are linear in any number of variables and those not linear in two independent variables, and also in regard to the function-theory of linear differential equations of the second order. space renders impossible anything further than the briefest account of many other matters; in particular, the theories of partial equations of higher than the first order, the function-theory of the singularities of ordinary equations not linear and the applications to differential geometry, are taken account of only in the bibliography. it is believed that on the whole the article will be more useful to the reader than if explanations of method had been further curtailed to include more facts. when we speak of a function without qualification, it is to be understood that in the immediate neighbourhood of a particular set x0, y0, ... of values of the independent variables x, y, ... of the function, at whatever point of the range of values for x, y, ... under consideration x0, y0, ... may be chosen, the function can be expressed as a series of positive integral powers of the differences x - x0, y -y0, ..., convergent when these are sufficiently small (see function: functions of complex variables). without this condition, which we express by saying that the function is developable about x0, y0, ..., many results provisionally stated in the transformation theories would be unmeaning or incorrect. if, then, we have a set of k functions, f1 ... fk of n independent variables x1 ... xn, we say that they are independent when n >= k and not every determinant of k rows and columns vanishes of the matrix of k rows and n columns whose r-th row has the constituents dfr/dx1, ... dfr/dxn; the justification being in the theorem, which we assume, that if the determinant involving, for instance, the first k columns be not zero for x1 = x1^0 ... xn = xn^0, and the functions be developable about this point, then from the equations f1 = c1, ... fk = ck we can express x1, ... xk by convergent power series in the differences x_k+1 - x_k+1^0, ... x_n - x_n^0, and so regard x1, ... xk as functions of the remaining variables. this we often express by saying that the equations f1 = c1, ... fk = ck can be solved for x1, ... xk. the explanation is given as a type of explanation often understood in what follows. ordinary equations of the first order. single homogeneous partial equation of the first order. proof of the existence of integrals. we may conveniently begin by stating the theorem: if each of the n functions [phi]1, ... [phi]n of the (n + 1) variables x1, ... x_nt be developable about the values x1^0, ... x_n^0t^0, the n differential equations of the form dx1/dt = [phi]1(tx1, ... xn) are satisfied by convergent power series x_r = x_r^0 + (t - t^0 ) a_r1 + (t - t0 )2a_r2 + ... reducing respectively to x1^0, ... xn^0 when t = t^0; and the only functions satisfying the equations and reducing respectively to x1^0, ... xn^0 when t = t^0, are those determined by continuation of these series. if the result of solving these n equations for x1^0, ... xn^0 be written in the form [omega]1(x1, ... xnt) = x1^0, ... [omega]n(x1, ... xnt) = xn^0, it is at once evident that the differential equation df/dt + [phi]1 df/dx1 + ... + [phi]n df/dxn = 0 possesses n integrals, namely, the functions [omega]1, ... [omega]n, which are developable about the values (x1^0 ... xn^0t^0) and reduce respectively to x1, ... xn when t = t^0. and in fact it has no other integrals so reducing. thus this equation also possesses a unique integral reducing when t = t^0 to an arbitrary function [psi](x1, ... xn), this integral being. [psi]([omega]1, ... [omega]n). conversely the existence of these _principal_ integrals [omega]1, ... [omega]n of the partial equation establishes the existence of the specified solutions of the ordinary equations dxi/dt = [phi]i. the following sketch of the proof of the existence of these principal integrals for the case n = 2 will show the character of more general investigations. put x for x - x^0, &c., and consider the equation a(xyt) df/dx + b(xyt) df/dy = df/dt, wherein the functions a, b are developable about x = 0, y = 0, t = 0; say a(xyt) = a0 + ta1 + t2a2/2! + ..., b(xyt) = b0 + tb1 + t2b2/2! + ..., so that ad/dx + bd/dy = [delta]0 + t[delta]1 + 1⁄2t2[delta]2 + ..., where [delta] = a_r d/dx + b_r d/dy. in order that f = p0 + tp1 + t2p2/2! + ... wherein p0, p1 ... are power series in x, y, should satisfy the equation, it is necessary, as we find by equating like terms, that p1 = [delta]0 p0, p2 = [delta]0 p1 + [delta]1 p0, &c. and in general p_s+1 = [delta]0 p_s + s1 [delta]1 p_s-1 + ... + [delta]_s p0, where s_r = (s!)/(r!) (s - r)! now compare with the given equation another equation a(xyt)df/dx + b(xyt)df/dy = df/dt, wherein each coefficient in the expansion of either a or b is real and positive, and not less than the absolute value of the corresponding coefficient in the expansion of a or b. in the second equation let us substitute a series f = p0 + tp1 + t2p2/2! + ..., wherein the coefficients in p0 are real and positive, and each not less than the absolute value of the corresponding coefficient in p0; then putting [delta]r = a_r d/dx + b_r d/dy we obtain necessary equations of the same form as before, namely, p1 = [delta]0 p0, p2= [delta]0 p1 + [delta]1 p0, ... and in general p_s+1 = [delta]0 p_s, + s1[delta]1 p_s-1 + ... + [delta]_s p0. these give for every coefficient in ps+1 an integral aggregate with real positive coefficients of the coefficients in p_s, p_s-1, ..., p0 and the coefficients in a and b; and they are the same aggregates as would be given by the previously obtained equations for the corresponding coefficients in p_s+1 in terms of the coefficients in ps, p_s-1, ..., p0 and the coefficients in a and b. hence as the coefficients in p0 and also in a, b are real and positive, it follows that the values obtained in succession for the coefficients in p1, p2, ... are real and positive; and further, taking account of the fact that the absolute value of a sum of terms is not greater than the sum of the absolute values of the terms, it follows, for each value of s, that every coefficient in p_s+1 is, in absolute value, not greater than the corresponding coefficient in p_s+1. thus if the series for f be convergent, the series for f will also be; and we are thus reduced to (1), specifying functions a, b with real positive coefficients, each in absolute value not less than the corresponding coefficient in a, b; (2) proving that the equation adf/dx + bdf/dy = df/dt possesses an integral p0 + tp1 + t2p2/2! + ... in which the coefficients in p0 are real and positive, and each not less than the absolute value of the corresponding coefficient in p0. if a, b be developable for x, y both in absolute value less than r and for t less in absolute value than r, and for such values a, b be both less in absolute value than the real positive constant m, it is not difficult to verify that we may take / x + y\-1 / t\-1 a = b = m( 1 - ----- ) ( 1 - - ), \ r / \ r/ and obtain _ _ | 4mr / x + y\-2 / t\-1 |1⁄2 f = r - (r - x - y) | 1 - ---(1 - ------) log (1 - - ) |, |_ r \ r / \ r/ _| and that this solves the problem when x, y, t are sufficiently small for the two cases p0 = x, p0 = y. one obvious application of the general theorem is to the proof of the existence of an integral of an ordinary linear differential equation given by the n equations dy/dx = y1, dy1/dx = y2, ..., dy_n-1/dx = p - p1 y_n-1 - ... - p_n y; but in fact any simultaneous system of ordinary equations is reducible to a system of the form dx1/dt = [phi](tx1, ... x_n). simultaneous linear partial equations. complete systems of linear partial equations. jacobian systems. suppose we have k homogeneous linear partial equations of the first order in n independent variables, the general equation being a_[sigma]1 df/dx1 + ... + a_[sigma]n df/dx_n = 0, where [sigma] = 1, ... k, and that we desire to know whether the equations have common solutions, and if so, how many. it is to be understood that the equations are linearly independent, which implies that k <= n and not every determinant of k rows and columns is identically zero in the matrix in which the i-th element of the [sigma]-th row is a[sigma]_i(i = 1, ... n, [sigma] = 1, ... k). denoting the left side of the [sigma]-th equation by p[sigma]f, it is clear that every common solution of the two equations p_[sigma]f = 0, p_[rho]f = 0, is also a solution of the equation p_[rho](p_[sigma]f), p_[sigma](p_[rho]f), we immediately find, however, that this is also a linear equation, namely, [sigma]h_i df/dx_i = 0 where h_i = p[rho]a[sigma]_i - p[sigma]a[rho]_i, and if it be not already contained among the given equations, or be linearly deducible from them, it may be added to them, as not introducing any additional limitation of the possibility of their having common solutions. proceeding thus with every pair of the original equations, and then with every pair of the possibly augmented system so obtained, and so on continually, we shall arrive at a system of equations, linearly independent of each other and therefore not more than n in number, such that the combination, in the way described, of every pair of them, leads to an equation which is linearly deducible from them. if the number of this so-called _complete system_ is n, the equations give df/dx1 = 0 ... df/dxn = 0, leading to the nugatory result f = a constant. suppose, then, the number of this system to be r < n; suppose, further, that from the matrix of the coefficients a determinant of r rows and columns not vanishing identically is that formed by the coefficients of the differential coefficients of f in regard to x1 ... x_r; also that the coefficients are all developable about the values x1 = x1^0, ... xn= xn^0, and that for these values the determinant just spoken of is not zero. then the main theorem is that the complete system of r equations, and therefore the originally given set of k equations, have in common n - r solutions, say [omega]r+1, ... [omega]n, which reduce respectively to x_r+1, ... x_n when in them for x1, ... x_r are respectively put x1^0, ... x_r^0; so that also the equations have in common a solution reducing when x1 = x1^0, ... x_r = x_r^0 to an arbitrary function [psi](x_r+1, ... x_n) which is developable about x_r+1^0, ... x_n^0, namely, this common solution is [psi]([omega]_r+1, ... [omega]_n). it is seen at once that this result is a generalization of the theorem for r = 1, and its proof is conveniently given by induction from that case. it can be verified without difficulty (1) that if from the r equations of the complete system we form r independent linear aggregates, with coefficients not necessarily constants, the new system is also a complete system; (2) that if in place of the independent variables x1, ... xn we introduce any other variables which are independent functions of the former, the new equations also form a complete system. it is convenient, then, from the complete system of r equations to form r new equations by solving separately for df/dx1, ..., df/dx_r; suppose the general equation of the new system to be q_[sigma]f = df/dx_[sigma] + c_[sigma],r+1 df/dx_r+1 + ... + c_[sigma]n df/dx_n = 0 ([sigma] = 1, ... r). then it is easily obvious that the equation q_[rho]q_[sigma]f - q_[sigma]q_[rho]f = 0 contains only the differential coefficients of f in regard to x_r+1 ... xn; as it is at most a linear function of q1f, ... qrf, it must be identically zero. so reduced the system is called a jacobian system. of this system q1f=0 has n - 1 principal solutions reducing respectively to x2, ... xn when x1 = x1^0, and its form shows that of these the first r - 1 are exactly x2 ... xr. let these n - 1 functions together with x1 be introduced as n new independent variables in all the r equations. since the first equation is satisfied by n - 1 of the new independent variables, it will contain no differential coefficients in regard to them, and will reduce therefore simply to df/dx1 = 0, expressing that any common solution of the r equations is a function only of the n - 1 remaining variables. thereby the investigation of the common solutions is reduced to the same problem for r - 1 equations in n - 1 variables. proceeding thus, we reach at length one equation in n - r + 1 variables, from which, by retracing the analysis, the proposition stated is seen to follow. system of total differential equations. the analogy with the case of one equation is, however, still closer. with the coefficients c_[sigma]j, of the equations q_[sigma]f = 0 in transposed array ([sigma] = 1, ... r, j = r + 1, ... n) we can put down the (n - r) equations, dx_j = c1_j dx1 + ... + c_rj dx_r, equivalent to the r(n - r) equations dx_j/dx_[sigma] = c_[sigma]r. that consistent with them we may be able to regard x_r+1, ... x_n as functions of x1, ... x_r, these being regarded as independent variables, it is clearly necessary that when we differentiate c_[sigma]j in regard to x_[rho] on this hypothesis the result should be the same as when we differentiate c[rho]j, in regard to x[sigma] on this hypothesis. the differential coefficient of a function f of x1, ... xn on this hypothesis, in regard to x_[rho]j is, however, df/dx_[rho] + c_[rho],r+1 df/dx_r+1 + ... + c_[rho]n df/dx_n, namely, is q_[rho]f. thus the consistence of the n - r total equations requires the conditions q_[rho]c_[sigma]j - q_[sigma]c_[rho]j = 0, which are, however, verified in virtue of q[rho](q[sigma][f]) - q_[sigma](q_[rho]f) = 0. and it can in fact be easily verified that if [omega]_r+1, ... [omega]_n be the principal solutions of the jacobian system, q_[sigma]f = 0, reducing respectively to x_r+1, ... xn when x1 = x1^0, ... x_r = x_r^0, and the equations [omega]_r+1 = x_r+1^0, ... [omega]_n = x_n^0 be solved for x_r+1, ... x_n to give x_j = [psi]_j(x1, ... x_r, x_r+1^0, ... x_n^0), these values solve the total equations and reduce respectively to x_r+1^0, ... x_n^0 when x1 = x1^0 ... x_r = x_r^0. and the total equations have no other solutions with these initial values. conversely, the existence of these solutions of the total equations can be deduced a priori and the theory of the jacobian system based upon them. the theory of such total equations, in general, finds its natural place under the heading _pfaffian expressions_, below. geometrical interpretation and solution. mayer's method of integration. a practical method of reducing the solution of the r equations of a jacobian system to that of a single equation in n - r + 1 variables may be explained in connexion with a geometrical interpretation which will perhaps be clearer in a particular case, say n = 3, r = 2. there is then only one total equation, say dz = adz + bdy; if we do not take account of the condition of integrability, which is in this case da/dy + bda/dz = db/dx + adb/dz, this equation may be regarded as defining through an arbitrary point (x0, y0, z0) of three-dimensioned space (about which a, b are developable) a plane, namely, z - z0 = a0(x - x0) + b0(y - y0), and therefore, through this arbitrary point [oo]2 directions, namely, all those in the plane. if now there be a surface z = [psi](x, y), satisfying dz = adz + bdy and passing through (x0, y0, z0), this plane will touch the surface, and the operations of passing along the surface from (x0, y0, z0) to (x0 + dx0, y0, z0 + dz0) and then to (x0 + dx0, y0 + dy0, z0 + d1z0), ought to lead to the same value of d^1z0 as do the operations of passing along the surface from (x0, y0, z0) to (x0, y0 + dy0, z0 + [delta]z0), and then to (x_ + dx_ , y_ + dy_ , z_ + [delta]1z_ ), 0 0 0 0 0 0 namely, [delta]1z0 ought to be equal to d1z0. but we find d1z0 = a0dx0 + b(x0 + dx0 , y0, z0 + a0dx0)dy0 = /db db \ a0dx0 + b0dy0 + dx0dy0( --- + a0--- ), \dx0 dz0/ and so at once reach the condition of integrability. if now we put x = x0 + t, y = y0 + mt, and regard m as constant, we shall in fact be considering the section of the surface by a fixed plane y - y0 = m(x - x0); along this section dz = dt(a + bm); if we then integrate the equation dx/dt = a + bm, where a, b are expressed as functions of m and t, with m kept constant, finding the solution which reduces to z0 for t = 0, and in the result again replace m by (y - y0)/(x - x0), we shall have the surface in question. in the general case the equations dx_j - c_1j dx1 + ... c_rj dx_r similarly determine through an arbitrary point x1^0, ... xn^0 a planar manifold of r dimensions in space of n dimensions, and when the conditions of integrability are satisfied, every direction in this manifold through this point is tangent to the manifold of r dimensions, expressed by [omega]_r+1 = x_r+1^0, ... [omega]_n = x_n^0, which satisfies the equations and passes through this point. if we put x1 = x1^0 = t, x2 = x2^0 = m2t, ... xr = xr^0 = mrt, and regard m2, ... mr as fixed, the (n-r) total equations take the form dx_j/dt = c_1j + m2c_2j + ... + m_rc_rj, and their integration is equivalent to that of the single partial equation n df/dt + [sigma](c_1j + m2c_2j + ... + m_rc_rj)df/dx_j = 0 j=r+1 in the n - r + 1 variables t, xr+1, ... xn. determining the solutions [omega]_r+1, ... [omega]_n which reduce to respectively x_r+1, ... x_n when t = 0, and substituting t = x1 - x1^0, m2 = (x2 - x2^0)/(x1 - x1^0), ... mr = (xr - xr^0)/(x1 - x1^0), we obtain the solutions of the original system of partial equations previously denoted by [omega]_r+1, ... [omega]_n. it is to be remarked, however, that the presence of the fixed parameters m2, ... mr in the single integration may frequently render it more difficult than if they were assigned numerical quantities. pfaffian expressions. we have above considered the integration of an equation dz = adz + bdy on the hypothesis that the condition da/dy + bda/dz = db/dz + adb/dz. it is natural to inquire what relations among x, y, z, if any, are implied by, or are consistent with, a differential relation adx + bdy + cdx = 0, when a, b, c are unrestricted functions of x, y, z. this problem leads to the consideration of the so-called _pfaffian expression_ adx + bdy + cdz. it can be shown (1) if each of the quantities db/dz - dc/dy, dc/dx - da/dz, da/dy - db/dz, which we shall denote respectively by u23, u31, u12, be identically zero, the expression is the differential of a function of x, y, z, equal to dt say; (2) that if the quantity au23 + bu31 + cu12 is identically zero, the expression is of the form udt, i.e. it can be made a perfect differential by multiplication by the factor 1/u; (3) that in general the expression is of the form dt + u1dt1. consider the matrix of four rows and three columns, in which the elements of the first row are a, b, c, and the elements of the (r+1)-th row, for r = 1, 2, 3, are the quantities u_r1, u_r2, u_r3, where u11 = u22 = u33 = 0. then it is easily seen that the cases (1), (2), (3) above correspond respectively to the cases when (1) every determinant of this matrix of two rows and columns is zero, (2) every determinant of three rows and columns is zero, (3) when no condition is assumed. this result can be generalized as follows: if a1, ... an be any functions of x1, ... xn, the so-called pfaffian expression a1dx1 + ... + a_ndx_n can be reduced to one or other of the two forms u1dt1 + ... + u_kdt_k, dt + u1dt1 + ... + u_k-1 dt_k-1, wherein t, u1 ..., t1, ... are independent functions of x1, ... xn, and k is such that in these two cases respectively 2k or 2k - 1 is the rank of a certain matrix of n + 1 rows and n columns, that is, the greatest number of rows and columns in a non-vanishing determinant of the matrix; the matrix is that whose first row is constituted by the quantities a1, ... an, whose s-th element in the (r+1)-th row is the quantity da_r/dx_s - da_s/dx_r. the proof of such a reduced form can be obtained from the two results: (1) if t be any given function of the 2m independent variables u1, ... um, t1, ... tm, the expression dt + u1 dt1 + ... + u_m dt_m can be put into the form u'1 dt'1 + ... + u'_mdt'_m. (2) if the quantities u1, ..., u1, t1, ... tm be connected by a relation, the expression n1dt1 + ... + umdtm can be put into the format dt' + u'1 dt'1 + ... + u'_m-1 dt'_m-1; and if the relation connecting u1, um, t1, ... tm be homogeneous in u1, ... um, then t' can be taken to be zero. these two results are deductions from the theory of _contact transformations_ (see below), and their demonstration requires, beside elementary algebraical considerations, only the theory of complete systems of linear homogeneous partial differential equations of the first order. when the existence of the reduced form of the pfaffian expression containing only independent quantities is thus once assured, the identification of the number k with that defined by the specified matrix may, with some difficulty, be made _a posteriori_. single linear pfaffian equation. in all cases of a single pfaffian equation we are thus led to consider what is implied by a relation dt - u1dt1 - ... - umdtm = 0, in which t, u1, ... um, t1 ..., tm are, except for this equation, independent variables. this is to be satisfied in virtue of one or several relations connecting the variables; these must involve relations connecting t, t1, ... tm only, and in one of these at least t must actually enter. we can then suppose that in one actual system of relations in virtue of which the pfaffian equation is satisfied, all the relations connecting t, t1 ... tm only are given by t = [psi](t_s+1 ... t_m), t1 = [psi]1(t_s+1 ... t_m), ... t_s = [psi]_s(t_s+1 ... t_m); so that the equation d[psi] - u1d[psi]1 - ... - u_s d[psi]_s - u_s+1 dt_s+1 - ... - u_m dt_m = 0 is identically true in regard to u1, ... um, t_s+1 ..., t_m; equating to zero the coefficients of the differentials of these variables, we thus obtain m - s relations of the form d[psi]/dt_j - u1 d[psi]1/dt_j - ... - u_s d[psi]_s/dt_j - u_j = 0; these m - s relations, with the previous s + 1 relations, constitute a set of m + 1 relations connecting the 2m + 1 variables in virtue of which the pfaffian equation is satisfied independently of the form of the functions [psi],[psi]1, ... [psi]s. there is clearly such a set for each of the values s = 0, s = 1, ..., s = m - 1, s = m. and for any value of s there may exist relations additional to the specified m + 1 relations, provided they do not involve any relation connecting t, t1, ... tm only, and are consistent with the m - s relations connecting u1, ... um. it is now evident that, essentially, the integration of a pfaffian equation a1dx1 + ... + a_n dx_n = 0, wherein a1, ... an are functions of x1, ... xn, is effected by the processes necessary to bring it to its reduced form, involving only independent variables. and it is easy to see that if we suppose this reduction to be carried out in all possible ways, there is no need to distinguish the classes of integrals corresponding to the various values of s; for it can be verified without difficulty that by putting t' = t - u1t1 - ... - u_s t_s, t'1 = u1, ... t'_s = u_s, u'1 = -t1, ..., u'_s = -t_s, t'_s+1 = t_s+1, ... t'_m = t_m, u'_s+1 = u_s+1, ... u'_m = u_m, the reduced equation becomes changed to dt' - u'1 dt'1 - ... - u'_m dt'_m = 0, and the general relations changed to t' = [psi](t'_s+l, ... t'_m) - t'1[psi]1(t'_s+1, ... t'_m) - ... -t'_s[psi]_s(t'_s+1, ... t'_m), = [phi], say, together with u'1 = d[phi]/dt'1, ..., u'm = d[phi]/dt'm, which contain only one relation connecting the variables t', t'1, ... t'm only. simultaneous pfaffian equations. this method for a single pfaffian equation can, strictly speaking, be generalized to a simultaneous system of (n - r) pfaffian equations dxj = c_1j dx1 + ... + c_rj dxr only in the case already treated, when this system is satisfied by regarding x_r+1, ... x_n as suitable functions of the independent variables x1, ... xr; in that case the integral manifolds are of r dimensions. when these are non-existent, there may be integral manifolds of higher dimensions; for if d[phi] = [phi]1 dx_r + ... + [phi]_r dx_r + [phi]_r+1(c_1,r+1 dx1 + ... + c_r,r+1 dx_r) + [phi]_r+2 ( ) + ... be identically zero, then [phi][sigma] + c[sigma]_,r+1 [phi]_r+1 + ... + c[sigma]_,n [phi]_n = 0, or [phi] satisfies the r partial differential equations previously associated with the total equations; when these are not a complete system, but included in a complete system of r - [mu] equations, having therefore n - r - [mu] independent integrals, the total equations are satisfied over a manifold of r + [mu] dimensions (see e. v. weber, _math. annal._ 1v. (1901), p. 386). contact transformations. it seems desirable to add here certain results, largely of algebraic character, which naturally arise in connexion with the theory of contact transformations. for any two functions of the 2n independent variables x1, ... xn, p1, ... pn we denote by ([phi][psi]) the sum of the n terms such as d[phi]d[psi]/dp_idx_i - d[psi]d[phi]/dp_idx_i. for two functions of the (2n + 1) independent variables z, x1, ... xn, p1, ... pn we denote by [phi][psi] the sum of the n terms such as d[phi] /d[psi] d[psi]\ d[psi] /d[phi] d[phi]\ ------( ------ + p_i------ ) - ------( ------ + p_i------ ). dpi \ dxi dz / dpi \ dxi dz / it can at once be verified that for any three functions [f[[phi][psi]]] + [[phi][psi]f]] + [[psi][f[phi]]] = df/dz [[phi][psi]] + d[phi]/dz [[psi]f] + d[psi]/dz [f[phi]], which when f, [phi],[psi] do not contain z becomes the identity (f([phi][psi])) + (phi([psi]f)) + ([psi](f[phi])) = 0. then, if x1, ... xn, p1, ... pn be such functions of x1, ... xn, p1 ... pn that p1 dx1 + ... + pn dxn is identically equal to p1dx1 + ... + pn dxn, it can be shown by elementary algebra, after equating coefficients of independent differentials, (1) that the functions x1, ... pn are independent functions of the 2n variables x1, ... pn, so that the equations x'i = xi, p'i = pi can be solved for x1, ... xn, p1, ... pn, and represent therefore a transformation, which we call a homogeneous contact transformation; (2) that the x1, ... xn are homogeneous functions of p1, ... pn of zero dimensions, the p1, ... pn are homogeneous functions of p1, ... pn of dimension one, and the 1⁄2n(n - 1) relations (xi xj) = 0 are verified. so also are the n2 relations (pi xi) = 1, (pi xj) = 0, (pi pj) = 0. conversely, if x1, ... xn be independent functions, each homogeneous of zero dimension in p1, ... pn satisfying the 1⁄2n(n - 1) relations (xi xj) = 0, then p1, ... pn can be uniquely determined, by solving linear algebraic equations, such that p1 dx1 + ... + pn dxn = p1 dx1 + ... + pn dxn. if now we put n + 1 for n, put z for x_n+1, z for x_n+1, qi for -pi/p_n+1, for i = 1, ... n, put qi for -p_i/p_n+1 and [sigma] for q_n+1/q_n+1, and then finally write p1, ... pn, p1, ... pn for q1, ... qn, q1, ... qn, we obtain the following results: if zx1 ... xn, p1, ... pn be functions of z, x1, ... xn, p1, ... pn, such that the expression dz - p1 dx1 - ... - pn dxn is identically equal to [sigma](dz - p1 dx1 - ... - pn dxn), and [sigma] not zero, then (1) the functions z, x1, ... xn, p1, ... pn are independent functions of z, x1, ... xn, p1, ... pn, so that the equations z' = z, x'i = xi, p'i = pi can be solved for z, x1, ... xn, p1, ... pn and determine a transformation which we call a (non-homogeneous) contact transformation; (2) the z, x1, ... xn verify the 1⁄2n(n + 1) identities [z xi] = 0, [xi xj] = 0. and the further identities [pi xi] = [sigma], [pi xj] = 0, [pi z] = [sigma]pi, [pi pj] = 0, dz dxi dpi [z[sigma]] = [sigma]-- - [sigma]2, [xi [sigma]] = [sigma]---, [pi [sigma]] = [sigma]--- dz dz dz are also verified. conversely, if z, x1, ... xn be independent functions satisfying the identities [z xi] = 0, [xi xj] = 0, then [sigma], other than zero, and p1, ... pn can be uniquely determined, by solution of algebraic equations, such that dz - p1 dx1 - ... - pn dxn = [sigma](dz - p1 dx1 - ... - p_n dx_n). finally, there is a particular case of great importance arising when [sigma] = 1, which gives the results: (1) if u, x1, ... xn, p1, ... pn be 2n + 1 functions of the 2n independent variables x1, ... xn, p1, ... pn, satisfying the identity du + p1 dx1 + ... + pn dxn = p1 dx1 + ... + p_n dx_n, then the 2n functions p1, ... pn, x1, ... xn are independent, and we have (xi xj) = 0, (xi u) = [delta]xi, (pi xi) = 1, (pi xj) = 0, (pi pj ) = 0, (pi u) + pi = [delta]pi, where [delta] denotes the operator p1d/dp1 + ... + pnd/dpn; (2) if x1, ... xn be independent functions of x1, ... xn, p1, ... pn, such that (xi xj) = 0, then u can be found by a quadrature, such that (xi u) = [delta]xi; and when xi, ... xn, u satisfy these 1⁄2n(n + 1) conditions, then p1, ... pn can be found, by solution of linear algebraic equations, to render true the identity du + p1 dx1 + ... + pn dxn = p1 dx1 + ... + pn dxn; (3) functions x1, ... xn, p1, ... pn can be found to satisfy this differential identity when u is an arbitrary given function of x1, ... xn, p1, ... pn; but this requires integrations. in order to see what integrations, it is only necessary to verify the statement that if u be an arbitrary given function of x1, ... xn, p1, ... pn, and, for r < n, x1, ... xr be independent functions of these variables, such that (x_[sigma] u) = [delta]x_[sigma], (x_[rho] x_[sigma]) = 0, for [rho], [sigma] = 1 ... r, then the r + 1 homogeneous linear partial differential equations of the first order (uf) + [delta]f = 0, (x[rho]f) = 0, form a complete system. it will be seen that the assumptions above made for the reduction of pfaffian expressions follow from the results here enunciated for contact transformations. partial differential equation of the first order. meaning of a solution of the equation. we pass on now to consider the solution of any partial differential equation of the first order; we attempt to explain certain ideas relatively to a single equation with any number of independent variables (in particular, an ordinary equation of the first order with one independent variable) by speaking of a single equation with two independent variables x, y, and one dependent variable z. it will be seen that we are naturally led to consider systems of such simultaneous equations, which we consider below. the central discovery of the transformation theory of the solution of an equation f(x, y, z, dz/dx, dz/dy) = 0 is that its solution can always be reduced to the solution of partial equations which are _linear_. for this, however, we must regard dz/dx, dz/dy, during the process of integration, not as the differential coefficients of a function z in regard to x and y, but as variables independent of x, y, z, the too great indefiniteness that might thus appear to be introduced being provided for in another way. we notice that if z = [psi](x, y) be a solution of the differential equation, then dz = dxd[psi]/dx + dyd[psi]/dy; thus if we denote the equation by f(x, y, z, p, q,) = 0, and prescribe the condition dz = pdx + qdy for every solution, any solution such as z = [psi](x, y) will necessarily be associated with the equations p = dz/dx, q = dz/dy, and z will satisfy the equation in its original form. we have previously seen (under _pfaffian expressions_) that if five variables x, y, z, p, q, otherwise independent, be subject to dz - pdx - qdy = 0, they must in fact be subject to at least three mutual relations. if we associate with a point (x, y, z) the plane z - z = p(x - x) + q(y - y) passing through it, where x, y, z are current co-ordinates, and call this association a surface-element; and if two consecutive elements of which the point(x + dx, y + dy, z + dz) of one lies on the plane of the other, for which, that is, the condition dz = pdx + qdy is satisfied, be said to be _connected,_ and an infinity of connected elements following one another continuously be called a _connectivity_, then our statement is that a connectivity consists of not more than [oo]2 elements, the whole number of elements (x, y, z, p, q) that are possible being called [oo]^5. the solution of an equation f(x, y, z, dz/dx, dz/dy) = 0 is then to be understood to mean finding in all possible ways, from the [oo]^4 elements (x, y, z, p, q) which satisfy f(x, y, z, p, q) = 0 a set of [oo]2 elements forming a connectivity; or, more analytically, finding in all possible ways two relations g = 0, h = 0 connecting x, y, z, p, q and independent of f = 0, so that the three relations together may involve dz = pdx + qdy. such a set of three relations may, for example, be of the form z = [psi](x, y), p = d[psi]/dx, q = d[psi]/dy; but it may also, as another case, involve two relations z = [psi](y), x = [psi]1(y) connecting x, y, z, the third relation being [psi]'(y) = p[psi]'1(y) + q, the connectivity consisting in that case, geometrically, of a curve in space taken with [oo]1 of its tangent planes; or, finally, a connectivity is constituted by a fixed point and all the planes passing through that point. this generalized view of the meaning of a solution of f = 0 is of advantage, moreover, in view of anomalies otherwise arising from special forms of the equation itself. for instance, we may include the case, sometimes arising when the equation to be solved is obtained by transformation from another equation, in which f does not contain either p or q. then the equation has [oo]2 solutions, each consisting of an arbitrary point of the surface f = 0 and all the [oo]2 planes passing through this point; it also has [oo]2 solutions, each consisting of a curve drawn on the surface f = 0 and all the tangent planes of this curve, the whole consisting of [oo]2 elements; finally, it has also an isolated (or singular) solution consisting of the points of the surface, each associated with the tangent plane of the surface thereat, also [oo]2 elements in all. or again, a linear equation f = pp + qq - r = 0, wherein p, q, r are functions of x, y, z only, has [oo]2 solutions, each consisting of one of the curves defined by dx/p = dy/q = dz/r taken with all the tangent planes of this curve; and the same equation has [oo]2 solutions, each consisting of the points of a surface containing [oo]1 of these curves and the tangent planes of this surface. and for the case of n variables there is similarly the possibility of n + 1 kinds of solution of an equation f(x1, ... xn, z, p1, ... pn) = 0; these can, however, by a simple contact transformation be reduced to one kind, in which there is only one relation z' = [psi](x'1, ... x'n) connecting the new variables x'1, ... x'n, z' (see under pfaffian expressions); just as in the case of the solution z = [psi](y), x = [psi]1(y), [psi]'(y) = p[psi]'1(y) + q of the equation pp + qq = r the transformation z' = z - px, x' = p, p' = -x, y' = y, q' = q gives the solution z' = [psi](y') + x'[psi]1(y'), p' = dz'/dx', q' = dz'/dy' of the transformed equation. these explanations take no account of the possibility of p and q being infinite; this can be dealt with by writing p = -u/w, q = -v/w, and considering homogeneous equations in u, v, w, with udx + vdy + wdz = 0 as the differential relation necessary for a connectivity; in practice we use the ideas associated with such a procedure more often without the appropriate notation. order of the ideas. in utilizing these general notions we shall first consider the theory of characteristic chains, initiated by cauchy, which shows well the nature of the relations implied by the given differential equation; the alternative ways of carrying out the necessary integrations are suggested by considering the method of jacobi and mayer, while a good summary is obtained by the formulation in terms of a pfaffian expression. characteristic chains. consider a solution of f = 0 expressed by the three independent equations f = 0, g = 0, h = 0. if it be a solution in which there is more than one relation connecting x, y, z, let new variables x', y', z', p', q' be introduced, as before explained under pfaffian expressions, in which z' is of the form z' = z - p1x1 - ... - p_s x_s (s = 1 or 2), so that the solution becomes of a form z' = [psi](x'y'), p' = d[psi]/dx', q' = d[psi]/dy', which then will identically satisfy the transformed equations f' = 0, g' = 0, h' = 0. the equation f' = 0, if x', y', z' be regarded as fixed, states that the plane z - z' = p'(x - x') + q'(y - y') is tangent to a certain cone whose vertex is (x', y', z'), the consecutive point (x' + dx', y' + dy', z' + dz') of the generator of contact being such that /df' /df' / / df' df'\ dx'/ -- = dy'/ -- = dz'/ ( p'--- + q' --- ). / dp' / dq' / \ dp' dq'/ passing in this direction on the surface z' = [psi](x', y') the tangent plane of the surface at this consecutive point is (p' + dp', q' + dq'), where, since f'(x', y', [psi], d[psi]/dx', d[psi]/dy') = 0 is identical, we have dx' (df'/dx' + p'df'/dz') + dp'df'/dp' = 0. thus the equations, which we shall call the characteristic equations, /df' /df' // df' df'\ // df' df'\ dx'/ --- = dy'/ --- = dz'/( p' --- + q'--- ) = dp'/( - --- - p'--- ) / dp' / dq' / \ dp' dq'/ / \ dx' dz'/ // df' df'\ = dq'/( - --- - q'--- ) / \ dy' dz'/ are satisfied along a connectivity of [oo]1 elements consisting of a curve on z' = [psi](x', y') and the tangent planes of the surface along this curve. the equation f' = 0, when p', q' are fixed, represents a curve in the plane z - z' = p'(x - x') + q'(y - y') passing through (x', y', z'); if (x' + [delta]x', y' + [delta]y', z' + [delta]z') be a consecutive point of this curve, we find at once /df' df'\ /df' df'\ [delta]x'( --- + p'--- ) + [delta]y'( --- + q'--- ) = 0; \dx' dz'/ \dy' dz'/ thus the equations above give [delta]x'dp' + [delta]y'dq' = 0, or the tangent line of the plane curve, is, on the surface z' = [psi](x', y'), in a direction conjugate to that of the generator of the cone. putting each of the fractions in the characteristic equations equal to dt, the equations enable us, starting from an arbitrary element x'0, y'0, z'0, p'0, q'0, about which all the quantities f', df'/dp', &c., occurring in the denominators, are developable, to define, from the differential equation f' = 0 alone, a connectivity of [oo]1 elements, which we call a _characteristic chain_; and it is remarkable that when we transform again to the original variables (x, y, z, p, q), the form of the differential equations for the chain is unaltered, so that they can be written down at once from the equation f = 0. thus we have proved that the characteristic chain starting from any ordinary element of any integral of this equation f = 0 consists only of elements belonging to this integral. for instance, if the equation do not contain p, q, the characteristic chain, starting from an arbitrary plane through an arbitrary point of the surface f = 0, consists of a pencil of planes whose axis is a tangent line of the surface f = 0. or if f = 0 be of the form pp + qq = r, the chain consists of a curve satisfying dx/p = dy/q = dz/r and a single infinity of tangent planes of this curve, determined by the tangent plane chosen at the initial point. in all cases there are [oo]3 characteristic chains, whose aggregate may therefore be expected to exhaust the [oo]^4 elements satisfying f = 0. complete integral constructed with characteristic chains. consider, in fact, a single infinity of connected elements each satisfying f = 0, say a chain connectivity t, consisting of elements specified by x0, y0, z0, p0, q0, which we suppose expressed as functions of a parameter u, so that u0 = dz0/du - p0dx0/du - q0dy0/du is everywhere zero on this chain; further, suppose that each of f, df/dp, ... , df/dx + pdf/dz is developable about each element of this chain t, and that t is _not_ a characteristic chain. then consider the aggregate of the characteristic chains issuing from all the elements of t. the [oo]2 elements, consisting of the aggregate of these characteristic chains, satisfy f = 0, provided the chain connectivity t consists of elements satisfying f = 0; for each characteristic chain satisfies df = 0. it can be shown that these chains are connected; in other words, that if x, y, z, p, q, be any element of one of these characteristic chains, not only is dz/dt - pdx/dt - qdy/dt = 0, as we know, but also u = dz/du - pdx/du - qdy/du is also zero. for we have du d /dz dx dy\ d /dz dx dy\ -- = --( -- - p-- - q-- ) - --( -- - p-- - q-- ) dt dt \du du du/ du \dt dt dt/ dp dx dp dx dq dy dq dy = -- -- - -- -- + -- -- - -- -- , du dt dt du du dt dt du which is equal to dp df dx /df df\ dq df dy /df df\ df -- -- + --( -- + p-- ) + -- -- + --( -- + q-- ) = - -- u. du dp du \dx dz/ du dq du \dy dz/ dz df as -- is a developable function of t, this, giving dz _ / / t df \ u = u_{0} exp( - | --dt ), \ _/t0 dz / shows that u is everywhere zero. thus integrals of f = 0 are obtainable by considering the aggregate of characteristic chains issuing from arbitrary chain connectivities t satisfying f = 0; and such connectivities t are, it is seen at once, determinable without integration. conversely, as such a chain connectivity t can be taken out from the elements of any given integral all possible integrals are obtainable in this way. for instance, an arbitrary curve in space, given by x0 = [theta](u), y0 = [phi](u), z0 = [psi](u), determines by the two equations f(x0, y0, z0, p0, q0) = 0, [psi]'(u) = p0[theta]'(u) + q0[phi]'(u), such a chain connectivity t, through which there passes a perfectly definite integral of the equation f = 0. by taking [oo]2 initial chain connectivities t, as for instance by taking the curves x0 = [theta], y0 = [phi], z0 = [psi] to be the [oo]2 curves upon an arbitrary surface, we thus obtain [oo]2 integrals, and so [oo]^4 elements satisfying f = 0. in general, if functions g, h, independent of f, be obtained, such that the equations f = 0, g = b, h = c represent an integral for all values of the constants b, c, these equations are said to constitute a _complete integral_. then [oo]^4 elements satisfying f = 0 are known, and in fact every other form of integral can be obtained without further integrations. operations necessary for integration of f = a. in the foregoing discussion of the differential equations of a characteristic chain, the denominators df/dp, ... may be supposed to be modified in form by means of f = 0 in any way conducive to a simple integration. in the immediately following explanation of ideas, however, we consider indifferently all equations f = constant; when a function of x, y, z, p, q is said to be zero, it is meant that this is so identically, not in virtue of f = 0; in other words, we consider the integration of f = a, where a is an arbitrary constant. in the theory of linear partial equations we have seen that the integration of the equations of the characteristic chains, from which, as has just been seen, that of the equation f = a follows at once, would be involved in completely integrating the single linear homogeneous partial differential equation of the first order [ff] = 0 where the notation is that explained above under contact transformations. one obvious integral is f = f. putting f = a, where a is arbitrary, and eliminating one of the independent variables, we can reduce this equation [ff] = 0 to one in four variables; and so on. calling, then, the determination of a single integral of a single homogeneous partial differential equation of the first order in n independent variables, _an operation of order_ n - 1, the characteristic chains, and therefore the most general integral of f = a, can be obtained by successive operations of orders 3, 2, 1. if, however, an integral of f = a be represented by f = a, g = b, h = c, where b and c are arbitrary constants, the expression of the fact that a characteristic chain of f = a satisfies dg = 0, gives [fg] = 0; similarly, [fh] = 0 and [gh] = 0, these three relations being identically true. conversely, suppose that an integral g, independent of f, has been obtained of the equation [ff] = 0, which is an operation of order three. then it follows from the identity [f[[phi][psi]]] + [[phi][[psi]f]] + [[psi][f[phi]]] = df/dz [[psi][phi]] + d[phi]/dz [psif] + d[psi]/dz [f[phi]] before remarked, by putting [phi] = f, [psi] = g, and then [ff] = a(f), [gf] = b(f), that ab(f) - ba(f) = df/dz b(f) - dg/dz a(f), so that the two linear equations [ff] = 0, [gf] = 0 form a complete system; as two integrals f, g are known, they have a common integral h, independent of f, g, determinable by an operation of order one only. the three functions f, g, h thus identically satisfy the relations [fg] = [gh] = [fh] = 0. the [oo]2 elements satisfying f = a, g = b, h = c, wherein a, b, c are assigned constants, can then be seen to constitute an integral of f = a. for the conditions that a characteristic chain of g = b issuing from an element satisfying f = a, g = b, h = c should consist only of elements satisfying these three equations are simply [fg] = 0, [gh] = 0. thus, starting from an arbitrary element of (f = a, g = b, h = c), we can single out a connectivity of elements of (f = a, g = b, h = c) forming a characteristic chain of g = b; then the aggregate of the characteristic chains of f = a issuing from the elements of this characteristic chain of g = b will be a connectivity consisting only of elements of (f = a, g = b, h = c), and will therefore constitute an integral of f = a; further, it will include all elements of (f = a, g = b, h = c). this result follows also from a theorem given under contact transformations, which shows, moreover, that though the characteristic chains of f = a are not determined by the three equations f = a, g = b, h = c, no further integration is now necessary to find them. by this theorem, since identically [fg] = [gh] = [fh] = 0, we can find, by the solution of linear algebraic equations only, a non-vanishing function [sigma] and two functions a, c, such that dg - adf - cdh = [sigma](dz - pdz - qdy); thus all the elements satisfying f = a, g = b, h = c, satisfy dz = pdx + qdy and constitute a connectivity, which is therefore an integral of f = a. while, further, from the associated theorems, f, g, h, a, c are independent functions and [fc] = 0. thus c may be taken to be the remaining integral independent of g, h, of the equation [ff] = 0, whereby the characteristic chains are entirely determined. the single equation f = 0 and pfaffian formulations. when we consider the particular equation f = 0, neglecting the case when neither p nor q enters, and supposing p to enter, we may express p from f = 0 in terms of x, y, z, q, and then eliminate it from all other equations. then instead of the equation [ff] = 0, we have, if f = 0 give p = [psi](x, y, z, q), the equation /df df\ d[psi] /df df\ /d[psi] d[psi]\ df [sigma]f = - ( -- + [psi] -- ) + ------ ( -- + q -- ) - ( ------ + q ------ ) -- = 0, \dx dz/ dq \dy dz/ \ dy dz / dq moreover obtainable by omitting the term in df/dp in [p-[psi], f] = 0. let x0, y0, z0, q0, be values about which the coefficients in this equation are developable, and let [zeta], [eta], [omega] be the principal solutions reducing respectively to z, y and q when x = x0. then the equations p = [psi], [zeta] = z0, [eta] = y0, [omega] = q0 represent a characteristic chain issuing from the element x0, y0, z0, [psi]0, q0; we have seen that the aggregate of such chains issuing from the elements of an arbitrary chain satisfying dz0 = p0dx0 - q0dy0 = 0 constitute an integral of the equation p = [psi]. let this arbitrary chain be taken so that x0 is constant; then the condition for initial values is only dz0 - q0dy0 = 0, and the elements of the integral constituted by the characteristic chains issuing therefrom satisfy d[zeta] - [omega]d[eta] = 0. hence this equation involves dz - [psi]dx - qdy = 0, or we have dz - [psi]dx - qdy = [sigma](d[zeta] - [omega]d[eta]), where [sigma] is not zero. conversely, the integration of p = [psi] is, essentially, the problem of writing the expression dz - [psi]dx - qdy in the form [sigma](d[zeta] - [omega]d[eta]), as must be possible (from what was said under _pfaffian expressions_). system of equations of the first order. to integrate a system of simultaneous equations of the first order x1 = a1, ... xr = ar in n independent variables x1, ... xn and one dependent variable z, we write p1 for dz/dx1, &c., and attempt to find n + 1 - r further functions z, x_r+1 ... xn, such that the equations z = a, xi = ai,(i = 1, ... n) involve dz - p1dx1 - ... - pndxn = 0. by an argument already given, the common integral, if existent, must be satisfied by the equations of the characteristic chains of any one equation xi = ai; thus each of the expressions [xi xj] must vanish in virtue of the equations expressing the integral, and we may without loss of generality assume that each of the corresponding 1⁄2r(r - 1) expressions formed from the r given differential equations vanishes in virtue of these equations. the determination of the remaining n + 1 - r functions may, as before, be made to depend on characteristic chains, which in this case, however, are manifolds of r dimensions obtained by integrating the equations [x1f] = 0, ... [xrf] = 0; or having obtained one integral of this system other than x1, ... xr, say xr+1, we may consider the system [x1f] = 0, ... [x_r+1 f] = 0, for which, again, we have a choice; and at any stage we may use mayer's method and reduce the simultaneous linear equations to one equation involving parameters; while if at any stage of the process we find some but not all of the integrals of the simultaneous system, they can be used to simplify the remaining work; this can only be clearly explained in connexion with the theory of so-called function groups for which we have no space. one result arising is that the simultaneous system p1 = [phi]1, ... pr = [phi]r, wherein p1, ... pr are not involved in [phi]1, ... [phi]r, if it satisfies the 1⁄2r(r - 1) relations [pi - [phi]i, pj - [phi]j] = 0, has a solution z = [psi](x1, ... xn), p1 = d[psi]/dx1, ... pn = d[psi]/dxn, reducing to an arbitrary function of x_r+1, ... xn only, when x1 = x1^0, ... xr = xr^0 under certain conditions as to developability; a generalization of the theorem for linear equations. the problem of integration of this system is, as before, to put dz - [phi]1dx1 - ... - [phi]_r dx_r - p_r+1 dx_r+1 - ... - p_n dx_n into the form [sigma](d[zeta] - [omega]_r+1 + d[xi]_r+1 - ... - [omega]_n d[xi]_n); and here [zeta], [xi]_r+1, ... [xi]_n, [omega]_r+1, ... [omega]_n may be taken, as before, to be principal integrals of a certain complete system of linear equations; those, namely, determining the characteristic chains. equations of dynamics. if l be a function of t and of the 2n quantities x1, ... xn, [.x]1, ... [.x]n, where [.x]i, denotes dxi/dt, &c., and if in the n equations d / dl \ dl --- (--------) = ---- dt \ dx_i / dx_i we put p_i = dl/d[.x]_i, and so express [.x]1 , ... [.x]_n in terms of t, x_i, ... x_n, p1, ... p_n, assuming that the determinant of the quantities d2l/dx_i d[.x]_j is not zero; if, further, h denote the function of t, x1, ... xn, p1, ... pn, numerically equal to p1[.x]1 + ... + pn[.x]n - l, it is easy to prove that dpi/dt = -dh/dxi, dxi/dt = dh/dp_i. these so-called _canonical_ equations form part of those for the characteristic chains of the single partial equation dz/dt + h(t, x1, ... xn, dz/dx1, ..., dz/dx_n) = 0, to which then the solution of the original equations for x1 ... xn can be reduced. it may be shown (1) that if z = [psi](t, x1, ... xn, c1, .. cn) + c be a complete integral of this equation, then pi = d[psi]/dx_i, d[psi]/dc_i = e_i are 2n equations giving the solution of the canonical equations referred to, where c1 ... cn and e1, ... en are arbitrary constants; (2) that if xi = xi(t, x^01, ... pn^0), pi=pi(t, x1^0, ... p^0n) be the principal solutions of the canonical equations for t = t^0, and [omega] denote the result of substituting these values in p1dh/dp1 + ... + pndh/dpn - h, and [omega] = [int] [t0 to t] [omega]dt, where, after integration, [omega] is to be expressed as a function of t, x1, ... xn, x1^0, ... xn^0, then z = [omega] + z^0 is a complete integral of the partial equation. application of theory of continuous groups to formal theories. a system of differential equations is said to allow a certain continuous group of transformations (see groups, theory of) when the introduction for the variables in the differential equations of the new variables given by the equations of the group leads, for all values of the parameters of the group, to the same differential equations in the new variables. it would be interesting to verify in examples that this is the case in at least the majority of the differential equations which are known to be integrable in finite terms. we give a theorem of very general application for the case of a simultaneous complete system of linear partial homogeneous differential equations of the first order, to the solution of which the various differential equations discussed have been reduced. it will be enough to consider whether the given differential equations allow the infinitesimal transformations of the group. it can be shown easily that sufficient conditions in order that a complete system [pi]1f = 0 ... [pi]kf = 0, in n independent variables, should allow the infinitesimal transformation pf = 0 are expressed by k equations [pi]_i pf - p[pi]_i f = [lambda]_i1 [pi]1f + ... + [lambda]_ik [pi]_kf. suppose now a complete system of n - r equations in n variables to allow a group of r infinitesimal transformations (p1f, ..., prf) which has an invariant subgroup of r - 1 parameters (p1f, ..., pr-1f), it being supposed that the n quantities [pi]1f, ..., [pi]_n-r f, p1 f, ..., p_r f are not connected by an identical linear equation (with coefficients even depending on the independent variables). then it can be shown that one solution of the complete system is determinable by a quadrature. for each of [pi]_i p_[sigma] f - p_[sigma] [pi]_i f is a linear function of [pi]1f, ..., [pi]_n-r f and the simultaneous system of independent equations [pi]1f = 0, ... [pi]_n-r f = 0, p1f = 0, ... p_r-1 f = 0 is therefore a complete system, allowing the infinitesimal transformation prf. this complete system of n - 1 equations has therefore one common solution [omega], and p_r([omega]) is a function of [omega]. by choosing [omega] suitably, we can then make pr([omega]) = 1. from this equation and the n - 1 equations [pi]_i[omega] = 0, p_[sigma][omega] = 0, we can determine [omega] by a quadrature only. hence can be deduced a much more general result, _that if the group of r parameters be integrable, the complete system can be entirety solved by quadratures_; it is only necessary to introduce the solution found by the first quadrature as an independent variable, whereby we obtain a complete system of n - r equations in n - 1 variables, subject to an integrable group of r - 1 parameters, and to continue this process. we give some examples of the application of the theorem. (1) if an equation of the first order y' = [psi](x, y) allow the infinitesimal transformation [xi]df/dx + [eta]df/dy, the integral curves [omega](x, y) = y°, wherein [omega](x, y) is the solution of df/dx + [psi](x, y) df/dy = 0 reducing to y for x = x°, are interchanged among themselves by the infinitesimal transformation, or [omega](x, y) can be chosen to make [xi]d[omega]/dx + [eta]d[omega]/dy = 1; this, with d[omega]/dx + [psi]d[omega]/dy = 0, determines [omega] as the integral of the complete differential (dy - [psi]dx)/([eta] - [psi][xi]). this result itself shows that every ordinary differential equation of the first order is subject to an infinite number of infinitesimal transformations. but every infinitesimal transformation [xi]df/dx + [eta]df/dy can by change of variables (after integration) be brought to the form df/dy, and all differential equations of the first order allowing this group can then be reduced to the form f(x, dy/dx) = 0. (2) in an ordinary equation of the second order y" = [psi](x, y, y'), equivalent to dy/dx = y1, dy1/dx = [psi](x, y, y1), if h, h1 be the solutions for y and y1 chosen to reduce to y^0 and y1° when x = x°, and the equations h = y, h1= y1 be equivalent to [omega] = y°, [omega]1 = y1°, then [omega], [omega]1 are the principal solutions of [pi]f = df/dx + y1df/dy + [psi]df/dy1 = 0. if the original equation allow an infinitesimal transformation whose first _extended_ form (see groups) is pf = [xi]df/dx + [eta]df/dy + [eta]1df/dy1, where [eta]1[delta]t is the increment of dy/dx when [xi][delta]t, [eta][delta]t are the increments of x, y, and is to be expressed in terms of x, y, y1, then each of p[omega] and p[omega]1 must be functions of [omega] and [omega]1, or the partial differential equation [pi]f must allow the group pf. thus by our general theorem, if the differential equation allow a group of two parameters (and such a group is always integrable), it can be solved by quadratures, our explanation sufficing, however, only provided the form [pi]f and the two infinitesimal transformations are not linearly connected. it can be shown, from the fact that [eta]1 is a quadratic polynomial in y1, that no differential equation of the second order can allow more than 8 really independent infinitesimal transformations, and that every homogeneous linear differential equation of the second order allows just 8, being in fact reducible to d2y/dx2 = 0. since every group of more than two parameters has subgroups of two parameters, a differential equation of the second order allowing a group of more than two parameters can, as a rule, be solved by quadratures. by transforming the group we see that if a differential equation of the second order allows a single infinitesimal transformation, it can be transformed to the form f(x, d[gamma]/dx, d2[gamma]/dx2); this is not the case for every differential equation of the second order. (3) for an ordinary differential equation of the third order, allowing an integrable group of three parameters whose infinitesimal transformations are not linearly connected with the partial equation to which the solution of the given ordinary equation is reducible, the similar result follows that it can be integrated by quadratures. but if the group of three parameters be simple, this result must be replaced by the statement that the integration is reducible to quadratures and that of a so-called riccati equation of the first order, of the form dy/dx = a + by + cy2, where a, b, c are functions of x. (4) similarly for the integration by quadratures of an ordinary equation yn = [psi](x, y, y1, ... yn-1) of any order. moreover, the group allowed by the equation may quite well consist of extended contact transformations. an important application is to the case where the differential equation is the resolvent equation defining the group of transformations or rationality group of another differential equation (see below); in particular, when the rationality group of an ordinary linear differential equation is integrable, the equation can be solved by quadratures. consideration of function theories of differential equations. following the practical and provisional division of theories of differential equations, to which we alluded at starting, into transformation theories and function theories, we pass now to give some account of the latter. these are both a necessary logical complement of the former, and the only remaining resource when the expedients of the former have been exhausted. while in the former investigations we have dealt only with values of the independent variables about which the functions are developable, the leading idea now becomes, as was long ago remarked by g. green, the consideration of the neighbourhood of the values of the variables for which this developable character ceases. beginning, as before, with existence theorems applicable for ordinary values of the variables, we are to consider the cases of failure of such theorems. a general existence theorem. when in a given set of differential equations the number of equations is greater than the number of dependent variables, the equations cannot be expected to have common solutions unless certain conditions of compatibility, obtainable by equating different forms of the same differential coefficients deducible from the equations, are satisfied. we have had examples in systems of linear equations, and in the case of a set of equations p1 = [phi]1, ..., pr = [phi]r. for the case when the number of equations is the same as that of dependent variables, the following is a general theorem which should be referred to: let there be r equations in r dependent variables z1, ... zr and n independent variables x1, ... xn; let the differential coefficient of z[sigma] of highest order which enters be of order h[sigma], and suppose d^h_[sigma] z_[sigma]/dx1^h_[sigma] to enter, so that the equations can be written d^h_[sigma] z_[sigma]/dx1^h_[sigma] = [phi]_[sigma], where in the general differential coefficient of z_[rho] which enters in [phi]_[sigma], say d^(k1 + ... + kn) z_[rho]/dx1^k1 ... dx_n^k_n, we have k1 < h_[rho] and k1 + ... + k_n <= h_[rho]. let a1, ... an, b1, ... br, and b[rho]_(k1 ... kn) be a set of values of x1, ... x_n, z1, ... z_r and of the differential coefficients entering in [phi]_[sigma] about which all the functions [phi]1, ... [phi]_r, are developable. corresponding to each dependent variable z_[sigma], we take now a set of h_[sigma] functions of x2, ... xn, say [phi][sigma], [phi][sigma]^(1), ..., [phi][sigma]^(h-1) arbitrary save that they must be developable about a2, a3, ... an, and such that for these values of x2, ... xn, the function [phi]_[rho] reduces to b_[rho], and the differential coefficient d^(k2 + ... + kn) [phi]_[rho]^(k1)/dx2^k2 ... dx_n^kn reduces to b^kn_(k1 ... kn). then the theorem is that there exists one, and only one, set of functions z1, ... z_r, of x2, ... x_n developable about a1, ... an satisfying the given differential equations, and such that for x1 = a1 we have z_[sigma] = [phi]_[sigma], dz_[sigma]/dx1 = [phi]_[sigma]^(1), ... d^(h_[sigma]-1) z_[sigma]/d^(h_[sigma]-1) x1 = [phi][sigma]^(h_[sigma]-1). and, moreover, if the arbitrary functions [phi]_[sigma], [phi]_[sigma]^(1) ... contain a certain number of arbitrary variables t1, ... tm, and be developable about the values t1°, ... tm° of these variables, the solutions z1, ... zr will contain t1, ... tm, and be developable about t1°, ... tm°. singular points of solutions. the proof of this theorem may be given by showing that if ordinary power series in x1 - -a1, ... xn - an, t1 - t1°, ... tm - tm° be substituted in the equations wherein in z[sigma] the coefficients of (x1 - a1)°, x1 - a1, ..., (x1 - a1)^(h_[sigma]-1) are the arbitrary functions [phi]_[sigma], [phi]_[sigma]^(1), ..., [phi]_[sigma]^h-1, divided respectively by 1, 1!, 2!, &c., then the differential equations determine uniquely all the other coefficients, and that the resulting series are convergent. we rely, in fact, upon the theory of monogenic analytical functions (see function), a function being determined entirely by its development in the neighbourhood of one set of values of the independent variables, from which all its other values arise by _continuation_; it being of course understood that the coefficients in the differential equations are to be continued at the same time. but it is to be remarked that there is no ground for believing, if this method of continuation be utilized, that the function is single-valued; we may quite well return to the same values of the independent variables with a different value of the function; belonging, as we say, to a different branch of the function; and there is even no reason for assuming that the number of branches is finite, or that different branches have the same singular points and regions of existence. moreover, and this is the most difficult consideration of all, all these circumstances may be dependent upon the values supposed given to the arbitrary constants of the integral; in other words, the singular points may be either _fixed_, being determined by the differential equations themselves, or they may be _movable_ with the variation of the arbitrary constants of integration. such difficulties arise even in establishing the reversion of an elliptic integral, in solving the equation /dx\2 ( -- ) = (x-a1)(x - a2)(x - a3)(x - a4); \ds/ about an ordinary value the right side is developable; if we put x - a1 = t12, the right side becomes developable about t1 = 0; if we put x = 1/t, the right side of the changed equation is developable about t = 0; it is quite easy to show that the integral reducing to a definite value x0 for a value s0 is obtainable by a series in integral powers; this, however, must be supplemented by showing that for no value of s does the value of x become entirely undetermined. linear differential equations with rational coefficients. these remarks will show the place of the theory now to be sketched of a particular class of ordinary linear homogeneous differential equations whose importance arises from the completeness and generality with which they can be discussed. we have seen that if in the equations dy/dx = y1, dy1/dx = y2, ..., dy_n-2/dx = y_n-1, dy_n-1/dx = a_n y + a_n-1 y1 + ... + a1 y_n-1, where a1, a2, ..., an are now to be taken to be rational functions of x, the value x = xo be one for which no one of these rational functions is infinite, and yo, yo1, ..., yo_n-1 be quite arbitrary finite values, then the equations are satisfied by y = you + yo1u1 + ... + yo_n-1 u_n-1, where u, u1, ..., un-1 are functions of x, independent of yo, ... yo_n-1, developable about x = xo; this value of y is such that for x = xo the functions y, y1 ... y_n-1 reduce respectively to yo, yo1, ... yo_n-1; it can be proved that the region of existence of these series extends within a circle centre xo and radius equal to the distance from xo of the nearest point at which one of a1, ... an becomes infinite. now consider a region enclosing xo and only one of the places, say [sigma], at which one of a1, ... an becomes infinite. when x is made to describe a closed curve in this region, including this point [sigma] in its interior, it may well happen that the continuations of the functions u, u1, ..., u_n-1 give, when we have returned to the point x, values v, v1, ..., v_n-1, so that the integral under consideration becomes changed to yo + yo1v1 + ... + yo_n-1 v_n-1. at xo let this branch and the corresponding values of y1, ... y_n-1 be [eta]o, [eta]o1, ... [eta]o_n-1; then, as there is only one series satisfying the equation and reducing to ([eta]o, [eta]o1, ... [eta]o_n-1) for x = xo and the coefficients in the differential equation are single-valued functions, we must have [eta]ou + [eta]o1u1 + ... + [eta]o_n-1 u_n-1 = yov + yo1v1 + ... + yo_n-1 v_n-1; as this holds for arbitrary values of yo ... yo_n-1, upon which u, ... u_n-1 and v, ... v_n-1 do not depend, it follows that each of v, ... v_n-1 is a linear function of u, ... u_n-1 with constant coefficients, say v_i = a_i1 u + ... + a_in u_n-1. then yov + ... + yo_n-1 v_n-1 = ([sigma]_i a_i1 y_io)u + ... + ([sigma]_i a_in yo_i)u_n-1; this is equal to [mu](you + ... + yo_n-1 u_n-1) if [sigma]_i a_ir yo_i = [mu]yo_r-1; eliminating yo ... yo_n-1 from these linear equations, we have a determinantal equation of order n for [mu]; let [mu]1 be one of its roots; determining the ratios of yo, y1o, ... yo_n-1 to satisfy the linear equations, we have thus proved that there exists an integral, h, of the equation, which when continued round the point [sigma] and back to the starting-point, becomes changed to h1 = [mu]1h. let now [xi] be the value of x at [sigma] and r1 one of the values of (1/2[pi]i) log [mu]1; consider the function (x - [xi])^r1 h; when x makes a circuit round x = [xi], this becomes changed to exp(-2[pi]ir1) (x - [xi])^-r1 [mu]h, that is, is unchanged; thus we may put h = (x - [xi])^r1 [phi]1, [phi]1 being a function single-valued for paths in the region considered described about [sigma], and therefore, by laurent's theorem (see function), capable of expression in the annular region about this point by a series of positive and negative integral powers of x - [xi], which in general may contain an infinite number of negative powers; there is, however, no reason to suppose r1 to be an integer, or even real. thus, if all the roots of the determinantal equation in [mu] are different, we obtain n integrals of the forms (x -[xi])^r1 phi1, ..., (x - [xi])^rn [phi]_n. in general we obtain as many integrals of this form as there are really different roots; and the problem arises to discover, in case a root be k times repeated, k - 1 equations of as simple a form as possible to replace the k - 1 equations of the form yo + ... + yo_n-1 v_n-1 = [mu](yo + ... + yo_n-1 u_n-1) which would have existed had the roots been different. the most natural method of obtaining a suggestion lies probably in remarking that if r2 = r1 + h, there is an integral [(x - [xi])^(r1 + h) [phi]2 - (x -[xi])^r1 [phi]1]/h, where the coefficients in [phi]2 are the same functions of r1 + h as are the coefficients in [phi]1 of r1; when h vanishes, this integral takes the form _ _ | d[phi]1 | (x - [xi])^r1 | ------- + [phi]1 log (x - [xi])|, |_ dr1 _| or say (x-[xi])^r1 [[phi]1 + [psi]1 log (x - [xi])]; denoting this by 2[pi]i[mu]1k, and (x-[xi])^r1 [phi]1 by h, a circuit of the point [xi] changes k into 1 k' = ----------- [e^(2[pi]ir1) (x - [xi])^r1 [psi]1 + e^(2[pi]ir1) (x - [xi])^r1 [phi]1 (2[pi]i + log(x - [xi]))] 2[pi]i[mu]1 = [mu]1k + h. a similar artifice suggests itself when three of the roots of the determinantal equation are the same, and so on. we are thus led to the result, which is justified by an examination of the algebraic conditions, that whatever may be the circumstances as to the roots of the determinantal equation, n integrals exist, breaking up into batches, the values of the constituents h1, h2, ... of a batch after circuit about x = [xi] being h1' = [mu]1h1, h2' = [mu]1h2 + h1, h3' = [mu]1h3 + h2, and so on. and this is found to lead to the forms (x - [xi])^r1 [phi]1, (x - [xi])^r1 [[psi]1 + [phi]1 log (x - [xi])], (x - [xi])^r1 [[chi]1 + [chi]2 log (x - [xi]) + [phi]1(log(x - [xi]))2], and so on. here each of [phi]1, [psi]1, [chi]1, [chi]2, ... is a series of positive and negative integral powers of x - [xi] in which the number of negative powers may be infinite. regular equations. it appears natural enough now to inquire whether, under proper conditions for the forms of the rational functions a1, ... an, it may be possible to ensure that in each of the series [phi]1, [psi]1, [chi]1, ... the number of negative powers shall be finite. herein lies, in fact, the limitation which experience has shown to be justified by the completeness of the results obtained. assuming n integrals in which in each of [phi]1, [psi]1, [chi]1 ... the number of negative powers is finite, there is a definite homogeneous linear differential equation having these integrals; this is found by forming it to have the form y'^n = (x - [xi])^-1 b1y'^(n-1) + (x - [xi])^-2 b2y'^(n-2) + ... +(x - [xi])^-n b_n y, where b1, ... bn are finite for x = [xi]. conversely, assume the equation to have this form. then on substituting a series of the form (x - [xi])^r [1 + a1(x - [xi]) + a2(x - [xi])2 + ... ] and equating the coefficients of like powers of x-[xi], it is found that r must be a root of an algebraic equation of order n; this equation, which we shall call the index equation, can be obtained at once by substituting for y only (x - [xi])^r and replacing each of b1, ... bn by their values at x = [xi]; arrange the roots r1, r2, ... of this equation so that the real part of ri is equal to, or greater than, the real part of r_i+1, and take r equal to r1; it is found that the coefficients a1, a2 ... are uniquely determinate, and that the series converges within a circle about x = [xi] which includes no other of the points at which the rational functions a1 ... an become infinite. we have thus a solution h1 = (x -[xi])^r1 [phi]1 of the differential equation. if we now substitute in the equation y = h1 f[eta]dx, it is found to reduce to an equation of order n - 1 for [eta] of the form [eta]'^(n-1) = (x - [xi])^-1 c1[eta]'^(n-2) + ... + (x-[xi])^(n-1) c_n-1 [eta], where c1, ... c_n-1 are not infinite at x = [xi]. to this equation precisely similar reasoning can then be applied; its index equation has in fact the roots r2 - r1 - 1, ... , rn - r1 - 1; if r2 - r1 be zero, the integral (x - [xi])^-1 [psi]1 of the [eta] equation will give an integral of the original equation containing log (x - [xi]); if r2 - r1 be an integer, and therefore a negative integer, the same will be true, unless in [psi]1 the term in (x - [xi])^(r1 - r2) be absent; if neither of these arise, the original equation will have an integral (x -[xi])^r2 [phi]2. the [eta] equation can now, by means of the one integral of it belonging to the index r2 - r1 - 1, be similarly reduced to one of order n - 2, and so on. the result will be that stated above. we shall say that an equation of the form in question is _regular_ about x = [xi]. fuchsian equations. equation of the second order. we may examine in this way the behaviour of the integrals at all the points at which any one of the rational functions a1 ... an becomes infinite; in general we must expect that beside these the value x = [oo] will be a singular point for the solutions of the differential equation. to test this we put x = 1/t throughout, and examine as before at t = 0. for instance, the ordinary linear equation with constant coefficients has no singular point for finite values of x; at x = [oo] it has a singular point and is not regular; or again, bessel's equation x2 + xy' + (x2 - n2)y = 0 is regular about x = 0, but not about x = [oo]. an equation regular at all the finite singularities and also at x = [oo] is called a fuchsian equation. we proceed to examine particularly the case of an equation of the second order y" + ay' + by = 0. putting x = 1/t, it becomes d2y/dt2 + (2t^-1 - at^-2)dy/dt + bt^-4 y = 0, which is not regular about t = 0 unless 2 - at^-1 and bt^-2, that is, unless ax and bx2 are finite at x =[oo]; which we thus assume; putting y = t^r(1 + a1t + ... ), we find for the index equation at x = [inifinity] the equation r(r - 1) + r(2 - ax)_0 + (bx2)_0 = 0. if there be finite singular points at [xi]1, ... [xi]m, where we assume m>1, the cases m = 0, m = 1 being easily dealt with, and if [phi](x) = (x - [xi]1) ... (x -[xi]m), we must have a.[phi](x) and b·[[phi](x)]2 finite for all finite values of x, equal say to the respective polynomials [psi](x) and [theta](x), of which by the conditions at x = [oo] the highest respective orders possible are m - 1 and 2(m - 1). the index equation at x = [xi]1 is r(r - 1) + r[psi]([xi]1)/[phi]'([xi]1) + [theta]([xi])1/[[phi]'([xi]1)]2 = 0, and if [alpha]1, [beta]1 be its roots, we have [alpha]1 + [beta]1 = 1 - [psi]([xi]1)/[phi]'([xi]1) and [alpha]1[beta]1 = [theta]([xi])1/[[phi]'([xi]1)]2. thus by an elementary theorem of algebra, the sum [sigma](1 - [alpha]i - [beta]i)/(x - [xi]i), extended to the m finite singular points, is equal to [psi](x)/[phi](x), and the sum [sigma](1 - [alpha]i - [beta]i) is equal to the ratio of the coefficients of the highest powers of x in [psi](x) and [phi](x), and therefore equal to 1 + [alpha] + [beta], where [alpha], [beta] are the indices at x = [oo]. further, if (x, 1)m-2 denote the integral part of the quotient [theta](x)/[phi](x), we have [sigma][alpha]_i[beta]_i[phi]'([xi]_i)/(x - [xi]_i) equal to -(x, 1)_m-2 + [theta](x)/[phi](x), and the coefficient of x^m-2 in (x, 1)_m-2 is [alpha][beta]. thus the differential equation has the form y" + y'[sigma](1 - [alpha]_i - [beta]_i)/(x - [xi]_i) + y[(x, 1)_m-2 + [sigma][alpha]_i[beta]_i[phi]'([xi]_i)/(x - [xi]_i)]/[phi](x) = 0. if, however, we make a change in the dependent variable, putting y = (x - [xi]1)^[alpha]1 ... (x - [xi]_m)^[alpha] m[eta], it is easy to see that the equation changes into one having the same singular points about each of which it is regular, and that the indices at x = [xi]_i become 0 and [beta]_i - [alpha]_i, which we shall denote by [lambda]i, for (x -[xi]_i)^[alpha]j can be developed in positive integral powers of x -[xi]_i about x = [xi]_i; by this transformation the indices at x = [oo] are changed to [alpha] + [alpha]1 + ... + [alpha]m, [beta] + [beta]1 + ... + [beta]m which we shall denote by [lambda], [mu]. if we suppose this change to have been introduced, and still denote the independent variable by y, the equation has the form y" + y'[sigma](1 - [lambda]_i)/(x - [xi]_i) + y(x, 1)_m-2/[phi](x) = 0, while [lambda] + [mu] + [lambda]1 + ... + [lambda]_m = m - 1. conversely, it is easy to verify that if [lambda][mu] be the coefficient of x^m-2 in (x, 1)_m-2, this equation has the specified singular points and indices whatever be the other coefficients in (x, 1)_m-2. hypergeometric equation. thus we see that (beside the cases m = 0, m = 1) the "fuchsian equation" of the second order with _two_ finite singular points is distinguished by the fact that it has a definite form when the singular points and the indices are assigned. in that case, putting (x - [xi]1)/(x - [xi]2) = t/(t - 1), the singular points are transformed to 0, 1, [oo], and, as is clear, without change of indices. still denoting the independent variable by x, the equation then has the form x(1 - x)y" + y'[1 - [lambda]1 - x(1 + [lambda] + [mu])] - [lambda][mu]y = 0, which is the ordinary hypergeometric equation. provided none of [lambda]1, [lambda]2, [lambda] - [mu] be zero or integral about x = 0, it has the solutions f([lambda], [mu], 1 - [lambda]1, x), x^[lambda]1 f([lambda] + [lambda]1, [mu] + [lambda]1, 1 + [lambda]1, x); about x = 1 it has the solutions f([lambda], [mu], 1 - [lambda]2, 1 - x), (1 - x)^[lambda]1 f([lambda] + [lambda]2, [mu] + [lambda]2, 1 + [lambda]2, 1 - x), where [lambda] + [mu] + [lambda]1 + [lambda]2 = 1; about x = [oo] it has the solutions x^-[lambda] f([lambda], [lambda] + [lambda]1, [lambda] - [mu] + 1, x^-1), x^-[mu] f([mu], [mu] + [lambda]1, [mu] - [lambda] + 1, x^-1), where f([alpha], [beta], [gamma], x) is the series [alpha][beta]x [alpha]([alpha] + 1)[beta]([beta] + 1)x2 1 + -------------- + ---------------------------------------- ..., [gamma] 1·2·[gamma]([gamma] + 1) which converges when |x| < 1, whatever [alpha], [beta], [gamma] may be, converges for all values of x for which |x| = 1 provided the real part of [gamma] - [alpha] - [beta] < 0 algebraically, and converges for all these values except x = 1 provided the real part of [gamma] - [alpha] -[beta] > -1 algebraically. in accordance with our general theory, logarithms are to be expected in the solution when one of [lambda]1, [lambda]2, [lambda] - [mu] is zero or integral. indeed when [lambda]1 is a negative integer, not zero, the second solution about x = 0 would contain vanishing factors in the denominators of its coefficients; in case [lambda] or [mu] be one of the positive integers 1, 2, ... (-[lambda]1), vanishing factors occur also in the numerators; and then, in fact, the second solution about x = 0 becomes x^[lambda]1 times an integral polynomial of degree (-[lambda]1) - [lambda] or of degree (-[lambda]1) - [mu]. but when [lambda]1 is a negative integer including zero, and neither [lambda] nor [mu] is one of the positive integers 1, 2 ... (-[lambda]1), the second solution about x = 0 involves a term having the factor log x. when [lambda]1 is a positive integer, not zero, the second solution about x = 0 persists as a solution, in accordance with the order of arrangement of the roots of the index equation in our theory; the first solution is then replaced by an integral polynomial of degree -[lambda] or -[mu]1, when [lambda] or [mu] is one of the negative integers 0, -1, -2, ..., 1 - [lambda]1, but otherwise contains a logarithm. similarly for the solutions about x = 1 or x = [oo]; it will be seen below how the results are deducible from those for x = 0. march of the integral. denote now the solutions about x = 0 by u1, u2; those about x = 1 by v1, v2; and those about x = [oo] by w1, w2; in the region (s0s1) common to the circles s0, s1 of radius 1 whose centres are the points x = 0, x = 1, all the first four are valid, and there exist equations u1 =av1 + bv2, u2 = cv1 + dv2 where a, b, c, d are constants; in the region (s1s) lying inside the circle s1 and outside the circle s0, those that are valid are v1, v2, w1, w2, and there exist equations v1 = pw1 + qw2, v2 = rw1 + tw2, where p, q, r, t are constants; thus considering any integral whose expression within the circle s0 is au1 + bu2, where a, b are constants, the same integral will be represented within the circle s1 by (aa + bc)v1 + (ab + bd)v2, and outside these circles will be represented by [(aa + bc)p + (ab + bd)r]w1 + [(aa + bc)q + (ab + bd)t]w2. a single-valued branch of such integral can be obtained by making a barrier in the plane joining [oo] to 0 and 1 to [oo]; for instance, by excluding the consideration of real negative values of x and of real positive values greater than 1, and defining the phase of x and x - 1 for real values between 0 and 1 as respectively 0 and [pi]. transformation of the equation into itself. we can form the fuchsian equation of the second order with three arbitrary singular points [xi]1, [xi]2, [xi]3, and no singular point at x = [oo], and with respective indices [alpha]1, [beta]1, [alpha]2, [beta]2, [alpha]3, [beta]3 such that [alpha]1 + [beta]1 + [alpha]2 + [beta]2 + [alpha]3 + [beta]3 = 1. this equation can then be transformed into the hypergeometric equation in 24 ways; for out of [xi]1, [xi]2, [xi]3 we can in six ways choose two, say [xi]1, [xi]2, which are to be transformed respectively into 0 and 1, by (x - [xi]1)/(x - [xi]2) = t(t - 1); and then there are four possible transformations of the dependent variable which will reduce one of the indices at t = 0 to zero and one of the indices at t = 1 also to zero, namely, we may reduce either [alpha]1 or [beta]1 at t = 0, and simultaneously either [alpha]2 or [beta]2 at t = 1. thus the hypergeometric equation itself can be transformed into itself in 24 ways, and from the expression f([lambda], [mu], 1 - [lambda]1, x) which satisfies it follow 23 other forms of solution; they involve four series in each of the arguments, x, x-1, 1/x, 1/(1-x), (x-1)/x, x/(x-1). five of the 23 solutions agree with the fundamental solutions already described about x = 0, x = 1, x = [oo]; and from the principles by which these were obtained it is immediately clear that the 24 forms are, in value, equal in fours. inversion. modular functions. the quarter periods k, k' of jacobi's theory of elliptic functions, of which k = [int] [0 to [pi]/2] (1 - h sin2[theta])^-1⁄2 d[theta], and k' is the same function of 1-h, can easily be proved to be the solutions of a hypergeometric equation of which h is the independent variable. when k, k' are regarded as defined in terms of h by the differential equation, the ratio k'/k is an infinitely many valued function of h. but it is remarkable that jacobi's own theory of theta functions leads to an expression for h in terms of k'/k (see function) in terms of single-valued functions. we may then attempt to investigate, in general, in what cases the independent variable x of a hypergeometric equation is a single-valued function of the ratio s of two independent integrals of the equation. the same inquiry is suggested by the problem of ascertaining in what cases the hypergeometric series f([alpha], [beta], [gamma], x) is the expansion of an algebraic (irrational) function of x. in order to explain the meaning of the question, suppose that the plane of x is divided along the real axis from -[oo] to 0 and from 1 to +[oo], and, supposing logarithms not to enter about x = 0, choose two quite definite integrals y1, y2 of the equation, say y1 = f([lambda], [mu], 1-[lambda]1, x), y2 = x^[lambda]1 f([lambda] + [lambda]1, [mu] + [lambda]1, 1 + [lambda]1, x), with the condition that the phase of x is zero when x is real and between 0 and 1. then the value of [sigma] = y2/y1 is definite for all values of x in the divided plane, [sigma] being a single-valued monogenic branch of an analytical function existing and without singularities all over this region. if, now, the values of [sigma] that so arise be plotted on to another plane, a value p + iq of [sigma] being represented by a point (p, q) of this [stigma]-plane, and the value of x from which it arose being mentally associated with this point of the [sigma]-plane, these points will fill a connected region therein, with a continuous boundary formed of four portions corresponding to the two sides of the two barriers of the x-plane. the question is then, firstly, whether the same value of s can arise for two different values of x, that is, whether the same point (p, q) of the [sigma]-plane can arise twice, or in other words, whether the region of the [sigma]-plane overlaps itself or not. supposing this is not so, a second part of the question presents itself. if in the x-plane the barrier joining -[oo] to 0 be momentarily removed, and x describe a small circle with centre at x = 0 starting from a point x = -h - ik, where h, k are small, real, and positive and coming back to this point, the original value s at this point will be changed to a value [sigma], which in the original case did not arise for this value of x, and possibly not at all. if, now, after restoring the barrier the values arising by continuation from [sigma] be similarly plotted on the s-plane, we shall again obtain a region which, while not overlapping itself, may quite possibly overlap the former region. in that case two values of x would arise for the same value or values of the quotient y2/y1, arising from two different branches of this quotient. we shall understand then, by the condition that x is to be a single-valued function of x, that the region in the [stimga]-plane corresponding to any branch is not to overlap itself, and that no two of the regions corresponding to the different branches are to overlap. now in describing the circle about x = 0 from x = -h - ik to -h + ik, where h is small and k evanescent, [stigma] = x^[lambda]1 f([lambda] + [lambda]1, [mu] + [lambda]1, 1 + [lambda]1, x)/f([lambda], [mu], 1 - [lambda]1, x) is changed to [sigma] = [stigma]e^(2[pi]i[lambda])1. thus the two portions of boundary of the s-region corresponding to the two sides of the barrier (-[oo], 0) meet (at [sigmaf] = 0 if the real part of [lambda]1 be positive) at an angle 2[pi]l1, where l1 is the absolute value of the real part of [lambda]1; the same is true for the [sigma]-region representing the branch [sigma]. the condition that the s-region shall not overlap itself requires, then, l1 = 1. but, further, we may form an infinite number of branches [sigma] = [stigma]e^(2[pi]i[lambda])1, [sigma]1 = e^(2[pi]i[lambda])1, ... in the same way, and the corresponding regions in the plane upon which y2/y1 is represented will have a common point and each have an angle 2[pi]l1; if neither overlaps the preceding, it will happen, if l1 is not zero, that at length one is reached overlapping the first, unless for some positive integer [alpha] we have 2[pi][alpha]l1 = 2[pi], in other words l1 = 1/a. if this be so, the branch [sigma]_a-1 = [stigma]e^(2[pi]ia[lambda])1 will be represented by a region having the angle at the common point common with the region for the branch [stigma]; but not altogether coinciding with this last region unless [lambda]1 be real, and therefore = ±1/a; then there is only a finite number, a, of branches obtainable in this way by crossing the barrier (-[oo], 0). in precisely the same way, if we had begun by taking the quotient [stigma]' = (x - 1)^[lambda]2 f([lambda] + [lambda]2, [mu] + [lambda]2, 1 + [lambda]2, 1 - x)/f([lambda], [mu], 1 - [lambda]2, 1 - x) of the two solutions about x = 1, we should have found that x is not a single-valued function of [stigma]' unless [lambda]2 is the inverse of an integer, or is zero; as [stigma]' is of the form (a[stigma] + b)/(c[stigma] + d), a, b, c, d constants, the same is true in our case; equally, by considering the integrals about x = [oo] we find, as a third condition necessary in order that x may be a single-valued function of [stigma], that [lambda] - [mu] must be the inverse of an integer or be zero. these three differences of the indices, namely, [lambda]1, [lambda]2, [lambda] - [mu], are the quantities which enter in the differential equation satisfied by x as a function of [stigma], which is easily found to be x111 32x211 - ---- + ------ = 1⁄2(h - h1 - h2)x^-1 (x - 1)^-1 + 1⁄2h1 x^-2 + 1⁄2h2(x - 1)^-2, x13 2x1^4 where x1 = dx/d[stigma], &c.; and h1 = 1 - y12, h2 = 1 - [lambda]22, h3 = 1 - ([lambda] - [mu])2. into the converse question whether the three conditions are sufficient to ensure (1) that the [stigma] region corresponding to any branch does not overlap itself, (2) that no two such regions overlap, we have no space to enter. the second question clearly requires the inquiry whether the group (that is, the monodromy group) of the differential equation is properly discontinuous. (see groups, theory of.) the foregoing account will give an idea of the nature of the function theories of differential equations; it appears essential not to exclude some explanation of a theory intimately related both to such theories and to transformation theories, which is a generalization of galois's theory of algebraic equations. we deal only with the application to homogeneous linear differential equations. rationality group of a linear equation. irreducibility of a rational equation. in general a function of variables x1, x2 ... is said to be rational when it can be formed from them and the integers 1, 2, 3, ... by a finite number of additions, subtractions, multiplications and divisions. we generalize this definition. assume that we have assigned a fundamental series of quantities and functions of x, in which x itself is included, such that all quantities formed by a finite number of additions, subtractions, multiplications, divisions _and differentiations in regard to x_, of the terms of this series, are themselves members of this series. then the quantities of this series, and only these, are called _rational_. by a rational function of quantities p, q, r, ... is meant a function formed from them and any of the fundamental rational quantities by a finite number of the five fundamental operations. thus it is a function which would be called, simply, rational if the fundamental series were widened by the addition to it of the quantities p, q, r, ... and those derivable from them by the five fundamental operations. a rational ordinary differential equation, with x as independent and y as dependent variable, is then one which equates to zero a rational function of y, the order k of the differential equation being that of the highest differential coefficient y^(k) which enters; only such equations are here discussed. such an equation p = 0 is called _irreducible_ when, firstly, being arranged as an integral polynomial in y^(k), this polynomial is not the product of other polynomials in y^(k) also of rational form; and, secondly, the equation has no solution satisfying also a rational equation of lower order. from this it follows that if an irreducible equation p = 0 have one solution satisfying another rational equation q = 0 of the same or higher order, then all the solutions of p = 0 also satisfy q = 0. for from the equation p = 0 we can by differentiation express y^(k+1), y^(k+2), ... in terms of x, y, y^(1), ... , y^(k), and so put the function q rationally in terms of these quantities only. it is sufficient, then, to prove the result when the equation q = 0 is of the same order as p = 0. let both the equations be arranged as integral polynomials in y^(k); their algebraic eliminant in regard to y^(k) must then vanish identically, for they are known to have one common solution not satisfying an equation of lower order; thus the equation p = 0 involves q = 0 for all solutions of p = 0. the variant function for a linear equation. now let y^(n) = [alpha]1y^(n-1) + ... + [alpha]_n y be a given rational homogeneous linear differential equation; let y1, ... yn be n particular functions of x, unconnected by any equation with constant coefficients of the form c1y1 + ... + cnyn = 0, all satisfying the differential equation; let [eta]1, ... [eta]n be linear functions of y1, ... yn, say [eta]i = a_i1 y1 + ... + a_in yn, where the constant coefficients aij have a non-vanishing determinant; write ([eta]) = a(y), these being the equations of a general linear homogeneous group whose transformations may be denoted by a, b, .... we desire to form a rational function [phi]([eta]), or say [phi](a(y)), of [eta]1, ... [eta], in which the [eta]2 constants aij shall all be essential, and not reduce effectively to a fewer number, as they would, for instance, if the y1, ... yn were connected by a linear equation with constant coefficients. such a function is in fact given, if the solutions y1, ... yn be developable in positive integral powers about x = a, by [phi]([eta]) = [eta]1 + (x - a)^n[eta]2 + ... + (x - a)^(n-1)n[eta]n. such a function, v, we call a _variant_. the resolvent eqution. then differentiating v in regard to x, and replacing [eta]i^(n) by its value a1[eta]^(n-1) + ... + an[eta], we can arrange dv/dx, and similarly each of d2/dx2 ... d^nv/dx^n, where n = n2, as a linear function of the n quantities [eta]1, ... [eta]n, ... [eta]1^(n-1), ... [eta]n^(n-1), and thence by elimination obtain a linear differential equation for v of order n with rational coefficients. this we denote by f = 0. further, each of [eta]1 ... [eta]n is expressible as a linear function of v, dv/dx, ... d^(n-1)v/dx^(n-1), with rational coefficients not involving any of the n2 coefficients a_ij, since otherwise v would satisfy a linear equation of order less than n, which is impossible, as it involves (linearly) the n2 arbitrary coefficients aij, which would not enter into the coefficients of the supposed equation. in particular, y1 ,.. yn are expressible rationally as linear functions of [omega], d[omega]/dx, ... d^(n-1)[omega]/dx^(n-1), where [omega] is the particular function [phi](y). any solution w of the equation f = 0 is derivable from functions [zeta]1, ... [zeta]n, which are linear functions of y1, ... yn, just as v was derived from [eta]1, ... [eta]n; but it does not follow that these functions [zeta]i, ... [zeta]n are obtained from y1, ... yn by a transformation of the linear group a, b, ... ; for it may happen that the determinant d([zeta]1, ... [zeta]n)/(dy1, ... yn) is zero. in that case [zeta]1, ... [zeta]n may be called a singular set, and w a singular solution; it satisfies an equation of lower than the n-th order. but every solution v, w, ordinary or singular, of the equation f = 0, is expressible rationally in terms of [omega], d[omega]/dx, ... d^(n-1)[omega]/dx^(n-1); we shall write, simply, v = r([omega]). consider now the rational irreducible equation of lowest order, not necessarily a linear equation, which is satisfied by [omega]; as y1, ... yn are particular functions, it may quite well be of order less than n; we call it the _resolvent equation_, suppose it of order p, and denote it by [gamma](v). upon it the whole theory turns. in the first place, as [gamma](v) = 0 is satisfied by the solution [omega] of f = 0, all the solutions of [gamma](v) are solutions f = 0, and are therefore rationally expressible by [omega]; any one may then be denoted by r([omega]). if this solution of f = 0 be not singular, it corresponds to a transformation a of the linear group (a, b, ...), effected upon y1, ... yn. the coefficients aij of this transformation follow from the expressions before mentioned for [eta]1 ... [eta]n in terms of v, dv/dx, d2v/dx2, ... by substituting v = r([omega]); thus they depend on the p arbitrary parameters which enter into the general expression for the integral of the equation [gamma](v) = 0. without going into further details, it is then clear enough that the resolvent equation, being irreducible and such that any solution is expressible rationally, with p parameters, in terms of the solution [omega], enables us to define a linear homogeneous group of transformations of y1 ... yn depending on p parameters; and every operation of this (continuous) group corresponds to a rational transformation of the solution of the resolvent equation. this is the group called the _rationality group_, or the _group of transformations_ of the original homogeneous linear differential equation. the group must not be confounded with a subgroup of itself, the _monodromy group_ of the equation, often called simply the group of the equation, which is a set of transformations, not depending on arbitrary variable parameters, arising for one particular fundamental set of solutions of the linear equation (see groups, theory of). the fundamental theorem in regard to the rationality group. the importance of the rationality group consists in three propositions. (1) any rational function of y1, ... yn which is unaltered in value by the transformations of the group can be written in rational form. (2) if any rational function be changed in form, becoming a rational function of y1, ... yn, a transformation of the group applied to its new form will leave its value unaltered. (3) any homogeneous linear transformation leaving unaltered the value of every rational function of y1, ... yn which has a rational value, belongs to the group. it follows from these that any group of linear homogeneous transformations having the properties (1) (2) is identical with the group in question. it is clear that with these properties the group must be of the greatest importance in attempting to discover what functions of x must be regarded as rational in order that the values of y1 ... yn may be expressed. and this is the problem of solving the equation from another point of view. literature.--([alpha]) _formal or transformation theories for equations of the first order_:--e. goursat, _lecons sur l'integration des equations aux derivees partielles du premier ordre_ (paris, 1891); e. v. weber, _vorlesungen uber das pfaff'sche problem und die theorie der partiellen differentialgleichungen erster ordnung_ (leipzig, 1900); s. lie und g. scheffers, _geometrie der beruhrungstransformationen_, bd. i. (leipzig, 1896); forsyth, _theory of differential equations, part i., exact equations and pfaff's problem_ (cambridge, 1890); s. lie, "allgemeine untersuchungen uber differentialgleichungen, die eine continuirliche endliche gruppe gestatten" (memoir), _mathem. annal._xxv. (1885), pp. 71-151; s. lie und g. scheffers, _vorlesungen uber differentialgleichungen mit bekannten infinitesimalen transformationen_ (leipzig, 1891). a very full bibliography is given in the book of e. v. weber referred to; those here named are perhaps sufficiently representative of modern works. of classical works may be named: jacobi, _vorlesungen uber dynamik_ (von a. clebsch, berlin, 1866); _werke, supplementband_; g monge, _application de l'analyse a la geometrie_ (par m. liouville, paris, 1850); j. l. lagrange, _lecons sur le calcul des fonctions_ (paris, 1806), and _theorie des fonctions analytiques_ (paris, prairial, an v); g. boole, _a treatise on differential equations_ (london, 1859); and _supplementary volume_ (london, 1865); darboux, _lecons sur la theorie generale des surfaces_, tt. i.-iv. (paris, 1887-1896); s. lie, _theorie der transformationsgruppen_ ii. (on contact transformations) (leipzig, 1890). ([beta]) _quantitative or function theories for linear equations_:--c. jordan, _cours d'analyse_, t. iii. (paris, 1896); e. picard, _traite d'analyse_, tt. ii. and iii. (paris, 1893, 1896); fuchs, _various memoirs, beginning with that in crelle's journal_, bd. lxvi. p. 121; riemann, _werke_, 2^r aufl. (1892); schlesinger, _handbuch der theorie der linearen differentialgleichungen_, bde. i.-ii. (leipzig, 1895-1898); heffter, _einleitung in die theorie der linearen differentialgleichungen mit einer unabhangigen variablen_ (leipzig, 1894); klein, _vorlesungen uber lineare differentialgleichungen der zweiten ordnung_ (autographed, gottingen, 1894); and _vorlesungen uber die hypergeometrische function_ (autographed, gottingen, 1894); forsyth, _theory of differential equations, linear equations_. ([gamma]) _rationality group (of linear differential equations)_:--picard, _traite d'analyse_, as above, t. iii.; vessiot, _annales de l'ecole normale_, serie iii. t. ix. p. 199 (memoir); s. lie, _transformationsgruppen_, as above, iii. a connected account is given in schlesinger, as above, bd. ii., erstes theil. ([delta]) _function theories of non-linear ordinary equations_:--painleve, _lecons sur la theorie analytique des equations differentielles_ (paris, 1897, autographed); forsyth, _theory of differential equations, part ii., ordinary equations not linear_ (two volumes, ii. and iii.) (cambridge, 1900); konigsberger, _lehrbuch der theorie der differentialgleichungen_ (leipzig, 1889); painleve, _lecons sur l'integration des equations differentielles de la mecanique et applications_ (paris, 1895). ([epsilon]) _formal theories of partial equations of the second and higher orders_:--e. goursat, _lecons sur l'integration des equations aux derivees partielles du second ordre_, tt. i. and ii. (paris, 1896, 1898); forsyth, _treatise on differential equations_ (london, 1889); and _phil. trans. roy. soc._ (a.), vol. cxci. (1898), pp. 1-86. ([zeta]) see also the six extensive articles in the second volume of the german _encyclopaedia of mathematics_. (h. f. ba.) difflugia (l. leclerc), a genus of lobose rhizopoda, characterized by a shell formed of sand granules cemented together; these are swallowed by the animal, and during the process of bud-fission they pass to the surface of the daughter-bud and are cemented there. _centropyxis_ (steia) and _lecqueureuxia_ (schlumberg) differ only in minor points. diffraction of light.--1. when light proceeding from a small source falls upon an opaque object, a shadow is cast upon a screen situated behind the obstacle, and this shadow is found to be bordered by alternations of brightness and darkness, known as "diffraction bands." the phenomena thus presented were described by grimaldi and by newton. subsequently t. young showed that in their formation interference plays an important part, but the complete explanation was reserved for a. j. fresnel. later investigations by fraunhofer, airy and others have greatly widened the field, and under the head of "diffraction" are now usually treated all the effects dependent upon the limitation of a beam of light, as well as those which arise from irregularities of any kind at surfaces through which it is transmitted, or at which it is reflected. 2. _shadows._--in the infancy of the undulatory theory the objection most frequently urged against it was the difficulty of explaining the very existence of shadows. thanks to fresnel and his followers, this department of optics is now precisely the one in which the theory has gained its greatest triumphs. the principle employed in these investigations is due to c. huygens, and may be thus formulated. if round the origin of waves an ideal closed surface be drawn, the whole action of the waves in the region beyond may be regarded as due to the motion continually propagated across the various elements of this surface. the wave motion due to any element of the surface is called a _secondary_ wave, and in estimating the total effect regard must be paid to the phases as well as the amplitudes of the components. it is usually convenient to choose as the surface of resolution a _wave-front_, i.e. a surface at which the primary vibrations are in one phase. any obscurity that may hang over huygens's principle is due mainly to the indefiniteness of thought and expression which we must be content to put up with if we wish to avoid pledging ourselves as to the character of the vibrations. in the application to sound, where we know what we are dealing with, the matter is simple enough in principle, although mathematical difficulties would often stand in the way of the calculations we might wish to make. the ideal surface of resolution may be there regarded as a flexible lamina; and we know that, if by forces locally applied every element of the lamina be made to move normally to itself exactly as the air at that place does, the external aerial motion is fully determined. by the principle of superposition the whole effect may be found by integration of the partial effects due to each element of the surface, the other elements remaining at rest. we will now consider in detail the important case in which uniform plane waves are resolved at a surface coincident with a wave-front (oq). we imagine a wave-front divided into elementary rings or zones--often named after huygens, but better after fresnel--by spheres described round p (the point at which the aggregate effect is to be estimated), the first sphere, touching the plane at o, with a radius equal to po, and the succeeding spheres with radii increasing at each step by 1⁄2[lambda]. there are thus marked out a series of circles, whose radii x are given by x2 + r2 = (r + 1⁄2n[lambda])2, or x2 = n[lambda]r nearly; so that the rings are at first of nearly equal area. now the effect upon p of each element of the plane is proportional to its area; but it depends also upon the distance from p, and possibly upon the inclination of the secondary ray to the direction of vibration and to the wave-front. o x q --------------------------- | / | / | / | / | / | / | / r| / | / | / | / | / | / | / | / | / p|/ fig. 1. the latter question can only be treated in connexion with the dynamical theory (see below, § 11); but under all ordinary circumstances the result is independent of the precise answer that may be given. all that it is necessary to assume is that the effects of the successive zones gradually diminish, whether from the increasing obliquity of the secondary ray or because (on account of the limitation of the region of integration) the zones become at last more and more incomplete. the component vibrations at p due to the successive zones are thus nearly equal in amplitude and opposite in phase (the phase of each corresponding to that of the infinitesimal circle midway between the boundaries), and the series which we have to sum is one in which the terms are alternately opposite in sign and, while at first nearly constant in numerical magnitude, gradually diminish to zero. in such a series each term may be regarded as very nearly indeed destroyed by the halves of its immediate neighbours, and thus the sum of the whole series is represented by half the first term, which stands over uncompensated. the question is thus reduced to that of finding the effect of the first zone, or central circle, of which the area is [pi][lambda]r. we have seen that the problem before us is independent of the law of the secondary wave as regards obliquity; but the result of the integration necessarily involves the law of the intensity and phase of a secondary wave as a function of r, the distance from the origin. and we may in fact, as was done by a. smith (_camb. math. journ._, 1843, 3, p. 46), determine the law of the secondary wave, by comparing the result of the integration with that obtained by supposing the primary wave to pass on to p without resolution. now as to the phase of the secondary wave, it might appear natural to suppose that it starts from any point q with the phase of the primary wave, so that on arrival at p, it is retarded by the amount corresponding to qp. but a little consideration will prove that in that case the series of secondary waves could not reconstitute the primary wave. for the aggregate effect of the secondary waves is the half of that of the first fresnel zone, and it is the central element only of that zone for which the distance to be travelled is equal to r. let us conceive the zone in question to be divided into infinitesimal rings of equal area. the effects due to each of these rings are equal in amplitude and of phase ranging uniformly over half a complete period. the phase of the resultant is midway between those of the extreme elements, that is to say, a quarter of a period behind that due to the element at the centre of the circle. it is accordingly necessary to suppose that the secondary waves start with a phase one-quarter of a period in advance of that of the primary wave at the surface of resolution. further, it is evident that account must be taken of the variation of phase in estimating the magnitude of the effect at p of the first zone. the middle element alone contributes without deduction; the effect of every other must be found by introduction of a resolving factor, equal to cos [theta], if [theta] represent the difference of phase between this element and the resultant. accordingly, the amplitude of the resultant will be less than if all its components had the same phase, in the ratio _ +1⁄2[pi] / | cos [theta]d[theta] : [pi], _/-1⁄2[pi] or 2 : [pi]. now 2 area /[pi] = 2[lambda]r; so that, in order to reconcile the amplitude of the primary wave (taken as unity) with the half effect of the first zone, the amplitude, at distance r, of the secondary wave emitted from the element of area ds must be taken to be ds/[lambda]r (1). by this expression, in conjunction with the quarter-period acceleration of phase, the law of the secondary wave is determined. that the amplitude of the secondary wave should vary as r^-1 was to be expected from considerations respecting energy; but the occurrence of the factor [lambda]^-1, and the acceleration of phase, have sometimes been regarded as mysterious. it may be well therefore to remember that precisely these laws apply to a secondary wave of sound, which can be investigated upon the strictest mechanical principles. the recomposition of the secondary waves may also be treated analytically. if the primary wave at o be cos kat, the effect of the secondary wave proceeding from the element ds at q is ds ds ------------- cos k(at - [rho] + 1⁄4[lambda]) = ------------- sin k(at - [rho]). [lambda][rho] [lambda][rho] if ds = 2[pi]xdx, we have for the whole effect _[oo] 2[pi] / sin k(at - [rho])x dx - -------- | ---------------------, [lambda] _/ 0 [rho] or, since xdx = [rho]d[rho], k = 2[pi]/[lambda], _[oo] _ _ / | |[oo] -k | sin k(at - [rho])d[rho] = | -cos k(at - [rho])| . _/r |_ _|r in order to obtain the effect of the primary wave, as retarded by traversing the distance r, viz. cos k(at - r), it is necessary to suppose that the integrated term vanishes at the upper limit. and it is important to notice that without some further understanding the integral is really ambiguous. according to the assumed law of the secondary wave, the result must actually depend upon the precise radius of the outer boundary of the region of integration, supposed to be exactly circular. this case is, however, at most very special and exceptional. we may usually suppose that a large number of the outer rings are incomplete, so that the integrated term at the upper limit may properly be taken to vanish. if a formal proof be desired, it may be obtained by introducing into the integral a factor such as e^-h[rho], in which h is ultimately made to diminish without limit. when the primary wave is plane, the area of the first fresnel zone is [pi][lambda]r, and, since the secondary waves vary as r^-1, the intensity is independent of r, as of course it should be. if, however, the primary wave be spherical, and of radius a at the wave-front of resolution, then we know that at a distance r further on the amplitude of the primary wave will be diminished in the ratio a:(r + a). this may be regarded as a consequence of the altered area of the first fresnel zone. for, if x be its radius, we have / {(r + 1⁄2[lambda])2 - x2} + \/ {a2 - x2} = r + a, so that x2 = [lambda]ar/(a + r) nearly. since the distance to be travelled by the secondary waves is still r, we see how the effect of the first zone, and therefore of the whole series is proportional to a/(a + r). in like manner may be treated other cases, such as that of a primary wave-front of unequal principal curvatures. the general explanation of the formation of shadows may also be conveniently based upon fresnel's zones. if the point under consideration be so far away from the geometrical shadow that a large number of the earlier zones are complete, then the illumination, determined sensibly by the first zone, is the same as if there were no obstruction at all. if, on the other hand, the point be well immersed in the geometrical shadow, the earlier zones are altogether missing, and, instead of a series of terms beginning with finite numerical magnitude and gradually diminishing to zero, we have now to deal with one of which the terms diminish to zero _at both ends_. the sum of such a series is very approximately zero, each term being neutralized by the halves of its immediate neighbours, which are of the opposite sign. the question of light or darkness then depends upon whether the series begins or ends abruptly. with few exceptions, abruptness can occur only in the presence of the first term, viz. when the secondary wave of least retardation is unobstructed, or when a _ray_ passes through the point under consideration. according to the undulatory theory the light cannot be regarded strictly as travelling along a ray; but the existence of an unobstructed ray implies that the system of fresnel's zones can be commenced, and, if a large number of these zones are fully developed and do not terminate abruptly, the illumination is unaffected by the neighbourhood of obstacles. intermediate cases in which a few zones only are formed belong especially to the province of diffraction. an interesting exception to the general rule that full brightness requires the existence of the first zone occurs when the obstacle assumes the form of a small circular disk parallel to the plane of the incident waves. in the earlier half of the 18th century r. delisle found that the centre of the circular shadow was occupied by a bright point of light, but the observation passed into oblivion until s. d. poisson brought forward as an objection to fresnel's theory that it required at the centre of a circular shadow a point as bright as if no obstacle were intervening. if we conceive the primary wave to be broken up at the plane of the disk, a system of fresnel's zones can be constructed which begin from the circumference; and the first zone external to the disk plays the part ordinarily taken by the centre of the entire system. the whole effect is the half of that of the first existing zone, and this is sensibly the same as if there were no obstruction. when light passes through a small circular or annular aperture, the illumination at any point along the axis depends upon the precise relation between the aperture and the distance from it at which the point is taken. if, as in the last paragraph, we imagine a system of zones to be drawn commencing from the inner circular boundary of the aperture, the question turns upon the manner in which the series terminates at the outer boundary. if the aperture be such as to fit exactly an integral number of zones, the aggregate effect may be regarded as the half of those due to the first and last zones. if the number of zones be even, the action of the first and last zones are antagonistic, and there is complete darkness at the point. if on the other hand the number of zones be odd, the effects conspire; and the illumination (proportional to the square of the amplitude) is four times as great as if there were no obstruction at all. the process of augmenting the resultant illumination at a particular point by stopping some of the secondary rays may be carried much further (soret, _pogg. ann._, 1875, 156, p. 99). by the aid of photography it is easy to prepare a plate, transparent where the zones of odd order fall, and opaque where those of even order fall. such a plate has the power of a condensing lens, and gives an illumination out of all proportion to what could be obtained without it. an even greater effect (fourfold) can be attained by providing that the stoppage of the light from the alternate zones is replaced by a phase-reversal without loss of amplitude. r. w. wood (_phil. mag._, 1898, 45, p 513) has succeeded in constructing zone plates upon this principle. in such experiments the narrowness of the zones renders necessary a pretty close approximation to the geometrical conditions. thus in the case of the circular disk, equidistant (r) from the source of light and from the screen upon which the shadow is observed, the width of the first exterior zone is given by dx = [lambda](2r)/4(2x), 2x being the diameter of the disk. if 2r = 1000 cm., 2x = 1 cm., [lambda] = 6 × 10^-5 cm., then dx = .0015 cm. hence, in order that this zone may be perfectly formed, there should be no error in the circumference of the order of .001 cm. (it is easy to see that the radius of the bright spot is of the same order of magnitude.) the experiment succeeds in a dark room of the length above mentioned, with a threepenny bit (supported by three threads) as obstacle, the origin of light being a small needle hole in a plate of tin, through which the sun's rays shine horizontally after reflection from an external mirror. in the absence of a heliostat it is more convenient to obtain a point of light with the aid of a lens of short focus. the amplitude of the light at any point in the axis, when plane waves are incident perpendicularly upon an annular aperture, is, as above, cos k(at - r1) - cos k(at - r2) = 2 sin kat sin k(r1 - r2), r2, r1 being the distances of the outer and inner boundaries from the point in question. it is scarcely necessary to remark that in all such cases the calculation applies in the first instance to homogeneous light, and that, in accordance with fourier's theorem, each homogeneous component of a mixture may be treated separately. when the original light is white, the presence of some components and the absence of others will usually give rise to coloured effects, variable with the precise circumstances of the case. [illustration: fig. 2.] although the matter can be fully treated only upon the basis of a dynamical theory, it is proper to point out at once that there is an element of assumption in the application of huygens's principle to the calculation of the effects produced by opaque screens of limited extent. properly applied, the principle could not fail; but, as may readily be proved in the case of sonorous waves, it is not in strictness sufficient to assume the expression for a secondary wave suitable when the primary wave is undisturbed, with mere limitation of the integration to the transparent parts of the screen. but, except perhaps in the case of very fine gratings, it is probable that the error thus caused is insignificant; for the incorrect estimation of the secondary waves will be limited to distances of a few wave-lengths only from the boundary of opaque and transparent parts. 3. _fraunhofer's diffraction phenomena._--a very general problem in diffraction is the investigation of the distribution of light over a screen upon which impinge divergent or convergent spherical waves after passage through various diffracting apertures. when the waves are convergent and the recipient screen is placed so as to contain the centre of convergency--the image of the original radiant point, the calculation assumes a less complicated form. this class of phenomena was investigated by j. von fraunhofer (upon principles laid down by fresnel), and are sometimes called after his name. we may conveniently commence with them on account of their simplicity and great importance in respect to the theory of optical instruments. if f be the radius of the spherical wave at the place of resolution, where the vibration is represented by cos kat, then at any point m (fig. 2) in the recipient screen the vibration due to an element ds of the wave-front is (§ 2) ds - ------------- sin k(at - [rho]), [lambda][rho] [rho] being the distance between m and the element ds. taking co-ordinates in the plane of the screen with the centre of the wave as origin, let us represent m by [xi], [eta], and p (where ds is situated) by x, y, z. then [rho]2 = (x - [xi])2 + (y - [eta])2 + z2, f2 = x2 + y2 + z2; so that [rho]2 = f2 - 2x[xi] - 2y[eta] + [xi]2 + [eta]2. in the applications with which we are concerned, [xi], [eta] are very small quantities; and we may take / x[xi] + y[eta]\ [rho] = f ( 1 - -------------- ). \ f2 / at the same time ds may be identified with dxdy, and in the denominator [rho] may be treated as constant and equal to f. thus the expression for the vibration at m becomes _ _ 1 / / / x[xi] + y[eta]\ - ------------- | | sin k ( at - f + -------------- ) dxdy (1); [lambda]2[f]2 _/_/ \ f / and for the intensity, represented by the square of the amplitude, _ _ _ _ 1 | / / x[xi] + y[eta] |2 i2 = ------------ | | | sin k -------------- dxdy | [lambda]2f2 |_ _/_/ f _| _ _ _ _ 1 | / / x[xi] + y[eta] |2 + ----------- | | | cos k -------------- dxdy | (2). [lambda]2f2 |_ _/_/ f _| this expression for the intensity becomes rigorously applicable when f is indefinitely great, so that ordinary optical aberration disappears. the incident waves are thus plane, and are limited to a plane aperture coincident with a wave-front. the integrals are then properly functions of the _direction_ in which the light is to be estimated. in experiment under ordinary circumstances it makes no difference whether the collecting lens is in front of or behind the diffracting aperture. it is usually most convenient to employ a telescope focused upon the radiant point, and to place the diffracting apertures immediately in front of the object-glass. what is seen through the eye-piece in any case is the same as would be depicted upon a screen in the focal plane. before proceeding to special cases it may be well to call attention to some general properties of the solution expressed by (2) (see bridge, _phil. mag._, 1858). if when the aperture is given, the wave-length (proportional to k^-1) varies, the composition of the integrals is unaltered, provided [xi] and [eta] are taken universely proportional to [lambda]. a diminution of [lambda] thus leads to a simple proportional shrinkage of the diffraction pattern, attended by an augmentation of brilliancy in proportion to [lambda]^-2. if the wave-length remains unchanged, similar effects are produced by an increase in the scale of the aperture. the linear dimension of the diffraction pattern is inversely as that of the aperture, and the brightness at corresponding points is as the _square_ of the area of aperture. if the aperture and wave-length increase in the same proportion, the size and shape of the diffraction pattern undergo no change. we will now apply the integrals (2) to the case of a rectangular aperture of width a parallel to x and of width b parallel to y. the limits of integration for x may thus be taken to be -1⁄2a and +1⁄2a, and for y to be -1⁄2b, +1⁄2b. we readily find (with substitution for k of 2[pi]/[lambda]) [pi]a[xi] [pi]b[eta] sin2 --------- sin2 ---------- a2b2 f[lambda] f[lambda] i2 = ----------- · ----------------- · --------------- (3), f2[lambda]2 [pi]2a2[xi]2 [pi]2b2[eta]2 ------------ ------------- f2[lambda]2 f2[lambda]2 as representing the distribution of light in the image of a mathematical point when the aperture is rectangular, as is often the case in spectroscopes. the second and third factors of (3) being each of the form sin2u/u2, we have to examine the character of this function. it vanishes when u = m[pi], m being any whole number other than zero. when u = 0, it takes the value unity. the maxima occur when u = tan u, (4), and then sin2u/u2 = cos2u (5). to calculate the roots of (5) we may assume u = (m + 1⁄2)[pi] - y = u - y, where y is a positive quantity which is small when u is large. substituting this, we find cot y = u - y, whence 1 / y y- \ y3 2y^5 17y^7 y = - ( 1 + - + -- + ... ) - -- ---- - -----. u \ u u2 / 3 15 315 this equation is to be solved by successive approximation. it will readily be found that 2 13 146 u = u - y = u - u^-1 - -- u^-3 - -- u^-5 - --- u^-7 - ... (6). 3 15 105 in the first quadrant there is no root after zero, since tan u > u, and in the second quadrant there is none because the signs of u and tan u are opposite. the first root after zero is thus in the third quadrant, corresponding to m = 1. even in this case the series converges sufficiently to give the value of the root with considerable accuracy, while for higher values of m it is all that could be desired. the actual values of u/[pi] (calculated in another manner by f. m. schwerd) are 1.4303, 2.4590, 3.4709, 4.4747, 5.4818, 6.4844, &c. since the maxima occur when u = (m + 1⁄2)[pi] nearly, the successive values are not very different from 4 4 4 ------, ------, -------, &c. 9[pi]2 25[pi] 49[pi]2 the application of these results to (3) shows that the field is brightest at the centre [xi] = 0, [eta] = 0, viz. at the geometrical image of the radiant point. it is traversed by dark lines whose equations are [xi] = mf[lambda]/a, [eta] = mf[lambda]/b. within the rectangle formed by pairs of consecutive dark lines, and not far from its centre, the brightness rises to a maximum; but these subsequent maxima are in all cases much inferior to the brightness at the centre of the entire pattern ([xi] = 0, [eta] = 0). by the principle of energy the illumination over the entire focal plane must be equal to that over the diffracting area; and thus, in accordance with the suppositions by which (3) was obtained, its value when integrated from [xi] = [oo] to [xi] = +[oo], and from [eta] = -[oo] to [eta] = +[oo] should be equal to ab. this integration, employed originally by p. kelland (_edin. trans._ 15, p. 315) to determine the absolute intensity of a secondary wave, may be at once effected by means of the known formula _+[oo] _+[oo] / sin2u / sin u | ----- du = | ----- du = [pi]. _/ u2 _/ u -[oo] -[oo] it will be observed that, while the total intensity is proportional to ab, the intensity at the focal point is proportional to a2b2. if the aperture be increased, not only is the total brightness over the focal plane increased with it, but there is also a concentration of the diffraction pattern. the form of (3) shows immediately that, if a and b be altered, the co-ordinates of any characteristic point in the pattern vary as a^-1 and b^-1. the contraction of the diffraction pattern with increase of aperture is of fundamental importance in connexion with the resolving power of optical instruments. according to common optics, where images are absolute, the diffraction pattern is supposed to be infinitely small, and two radiant points, however near together, form separated images. this is tantamount to an assumption that [lambda] is infinitely small. the actual finiteness of [lambda] imposes a limit upon the separating or resolving power of an optical instrument. this indefiniteness of images is sometimes said to be due to diffraction by the edge of the aperture, and proposals have even been made for curing it by causing the transition between the interrupted and transmitted parts of the primary wave to be less abrupt. such a view of the matter is altogether misleading. what requires explanation is not the imperfection of actual images so much as the possibility of their being as good as we find them. at the focal point ([xi] = 0, [eta] = 0) all the secondary waves agree in phase, and the intensity is easily expressed, whatever be the form of the aperture. from the general formula (2), if a be the _area_ of aperture, i02 = a2/[lambda]2f2 (7). the formation of a sharp image of the radiant point requires that the illumination become insignificant when [xi], [eta] attain small values, and this insignificance can only arise as a consequence of discrepancies of phase among the secondary waves from various parts of the aperture. so long as there is no sensible discrepancy of phase there can be no sensible diminution of brightness as compared with that to be found at the focal point itself. we may go further, and lay it down that there can be no considerable loss of brightness until the difference of phase of the waves proceeding from the nearest and farthest parts of the aperture amounts to 1⁄4[lambda]. when the difference of phase amounts to [lambda], we may expect the resultant illumination to be very much reduced. in the particular case of a rectangular aperture the course of things can be readily followed, especially if we conceive f to be infinite. in the direction (suppose horizontal) for which [eta] = 0, [xi]/f = sin [theta], the phases of the secondary waves range over a complete period when sin [theta] = [lambda]/a, and, since all parts of the horizontal aperture are equally effective, there is in this direction a complete compensation and consequent absence of illumination. when sin [theta] = 3/2[lambda]/a, the phases range one and a half periods, and there is revival of illumination. we may compare the brightness with that in the direction [theta] = 0. the phase of the resultant amplitude is the same as that due to the central secondary wave, and the discrepancies of phase among the components reduce the amplitude in the proportion _+3/2[pi] 1 / ----- | cos [phi] d[phi]: 1, 3[pi] _/-3/2[pi] or -2/3[pi]:1; so that the brightness in this direction is 4/9[pi]2 of the maximum at [theta] = 0. in like manner we may find the illumination in any other direction, and it is obvious that it vanishes when sin [theta] is any multiple of [lamba]/a. the reason of the augmentation of resolving power with aperture will now be evident. the larger the aperture the smaller are the angles through which it is necessary to deviate from the principal direction in order to bring in specified discrepancies of phase--the more concentrated is the image. in many cases the subject of examination is a luminous line of uniform intensity, the various points of which are to be treated as independent sources of light. if the image of the line be [xi] = 0, the intensity at any point [xi], [eta] of the diffraction pattern may be represented by [pi]a[xi] _+[oo] sin2--------- / a2b [lambda]f | i2d[eta] = --------- ------------- (8), _/ [lambda]f [pi]2a2[xi]2 -[oo] ------------ [lambda]2f2 the same law as obtains for a luminous point when horizontal directions are alone considered. the definition of a fine vertical line, and consequently the resolving power for contiguous vertical lines, is thus _independent of the vertical aperture of the instrument_, a law of great importance in the theory of the spectroscope. the distribution of illumination in the image of a luminous line is shown by the curve abc (fig. 3), representing the value of the function sin2u/u2 from u = 0 to u = 2[pi]. the part corresponding to negative values of u is similar, oa being a line of symmetry. [illustration: fig. 3.] let us now consider the distribution of brightness in the image of a double line whose components are of equal strength, and at such an angular interval that the central line in the image of one coincides with the first zero of brightness in the image of the other. in fig. 3 the curve of brightness for one component is abc, and for the other oa'c'; and the curve representing half the combined brightnesses is e'be. the brightness (corresponding to b) midway between the two central points aa' is .8106 of the brightness at the central points themselves. we may consider this to be about the limit of closeness at which there could be any decided appearance of resolution, though doubtless an observer accustomed to his instrument would recognize the duplicity with certainty. the obliquity, corresponding to u = [pi], is such that the phases of the secondary waves range over a complete period, i.e. such that the projection of the horizontal aperture upon this direction is one wave-length. we conclude that a _double line cannot be fairly resolved unless its components subtend an angle exceeding that subtended by the wave-length of light at a distance equal to the horizontal aperture_. this rule is convenient on account of its simplicity; and it is sufficiently accurate in view of the necessary uncertainty as to what exactly is meant by resolution. if the angular interval between the components of a double line be half as great again as that supposed in the figure, the brightness midway between is .1802 as against 1.0450 at the central lines of each image. such a falling off in the middle must be more than sufficient for resolution. if the angle subtended by the components of a double line be twice that subtended by the wave-length at a distance equal to the horizontal aperture, the central bands are just clear of one another, and there is a line of absolute blackness in the middle of the combined images. the resolving power of a telescope with circular or rectangular aperture is easily investigated experimentally. the best object for examination is a grating of fine wires, about fifty to the inch, backed by a sodium flame. the object-glass is provided with diaphragms pierced with round holes or slits. one of these, of width equal, say, to one-tenth of an inch, is inserted in front of the object-glass, and the telescope, carefully focused all the while, is drawn gradually back from the grating until the lines are no longer seen. from a measurement of the maximum distance the least angle between consecutive lines consistent with resolution may be deduced, and a comparison made with the rule stated above. merely to show the dependence of resolving power on aperture it is not necessary to use a telescope at all. it is sufficient to look at wire gauze backed by the sky or by a flame, through a piece of blackened cardboard, pierced by a needle and held close to the eye. by varying the distance the point is easily found at which resolution ceases; and the observation is as sharp as with a telescope. the function of the telescope is in fact to allow the use of a wider, and therefore more easily measurable, aperture. an interesting modification of the experiment may be made by using light of various wave-lengths. since the limitation of the width of the central band in the image of a luminous line depends upon discrepancies of phase among the secondary waves, and since the discrepancy is greatest for the waves which come from the edges of the aperture, the question arises how far the operation of the central parts of the aperture is advantageous. if we imagine the aperture reduced to two equal narrow slits bordering its edges, compensation will evidently be complete when the projection on an oblique direction is equal to 1⁄2[lambda], instead of [lambda] as for the complete aperture. by this procedure the width of the central band in the diffraction pattern is halved, and so far an advantage is attained. but, as will be evident, the bright bands bordering the central band are now not inferior to it in brightness; in fact, a band similar to the central band is reproduced an indefinite number of times, so long as there is no sensible discrepancy of phase in the secondary waves proceeding from the various parts of the _same_ slit. under these circumstances the narrowing of the band is paid for at a ruinous price, and the arrangement must be condemned altogether. a more moderate suppression of the central parts is, however, sometimes advantageous. theory and experiment alike prove that a double line, of which the components are equally strong, is better resolved when, for example, one-sixth of the horizontal aperture is blocked off by a central screen; or the rays quite at the centre may be allowed to pass, while others a little farther removed are blocked off. stops, each occupying one-eighth of the width, and with centres situated at the points of trisection, answer well the required purpose. it has already been suggested that the principle of energy requires that the general expression for i2 in (2) when integrated over the whole of the plane [xi], [eta] should be equal to a, where a is the area of the aperture. a general analytical verification has been given by sir g. g. stokes (_edin. trans._, 1853, 20, p. 317). analytically expressed-- _ _+[oo] _ _ / / / / | | i2 d[xi]d[eta] = | | dxdy = a (9). _/_/-[oo] _/_/ we have seen that i02 (the intensity at the focal point) was equal to a2/[lambda]2f2. if a' be the area over which the intensity must be i02 in order to give the actual total intensity in accordance with _ _+[oo] / / a'i02 = | | i2 d[xi]d[eta], _/_/-[oo] the relation between a and a' is aa' = [lambda]2f2. since a' is in some sense the area of the diffraction pattern, it may be considered to be a rough criterion of the definition, and we infer that the definition of a point depends principally upon the area of the aperture, and only in a very secondary degree upon the shape when the area is maintained constant. 4. _theory of circular aperture._--we will now consider the important case where the form of the aperture is circular. writing for brevity k[xi]/f = p, k[eta]/f = q, (1), we have for the general expression (§ 11) of the intensity [lambda]2f2i2 = s2 + c2 (2), where _ _ / / s = | | sin(px + qy)dx dy, (3), _/_/ _ _ / / c = | | cos(px + qy)dx dy, (4). _/_/ when, as in the application to rectangular or circular apertures, the form is symmetrical with respect to the axes both of x and y, s = 0, and c reduces to _ _ / / c = | | cos px cos qy dx dy, (5). _/_/ in the case of the circular aperture the distribution of light is of course symmetrical with respect to the focal point p = 0, q = 0; and c is a function of p and q only through [sqrt](p2 + q2). it is thus sufficient to determine the intensity along the axis of p. putting q = 0, we get _ _ _+r / / / / c = | | cos px dx dy = 2 | cos px \/(r2 - x2) dx, _/_/ _/-r r being the radius of the aperture. this integral is the bessel's function of order unity, defined by _[pi] z / j1(z) = ---- | cos(z cos [phi]) sin2 [phi] d[phi] (6). [pi] _/0 thus, if x = r cos [phi], 2j1(pr) c = [pi]2r ------- (7); pr and the illumination at distance r from the focal point is / 2[pi]rr \ 4j12( --------- ) [pi]2r^4 \f[lambda]/ i2 = ----------- · ----------------- (8). [lambda]2f2 / 2[pi]rr \2 ( --------- ) \f[lambda]/ the ascending series for j1(z), used by sir g. b. airy (_camb. trans._, 1834) in his original investigation of the diffraction of a circular object-glass, and readily obtained from (6), is z z3 z^5 z^7 j1(z) = - - ---- + ------- - ---------- + ... (9). 2 22·4 22·42·6 22·42·62·8 when z is great, we may employ the semi-convergent series _ / / 2 \ | 3·5·1 /1\2 j1(z) = / ( ----- ) sin (z - 1⁄4[pi]) |1 + ------ ( - ) \/ \[pi]z/ |_ 8·16 \z/ _ 3·5·7·9·1·3·5 /1\^4 | - ------------- ( - ) + ... | 8·16·24·32 \z/ _| _ / / 2 \ | 3 1 3·5·7·1·3 /1\ 3 + / ( ----- ) cos (z - 1⁄4[pi]) | - · - - --------- ( - ) \/ \[pi]z/ |_8 z 8·16·24 \z/ _ 3·5·7·9·11·1·3·5·7 /1\^5 | + ------------------ ( - ) - ... | ... (10). 8·16·24·32·40 \z/ _| a table of the values of 2z^-1j1(z) has been given by e. c. j. lommel (_schlomilch_, 1870, 15, p. 166), to whom is due the first systematic application of bessel's functions to the diffraction integrals. the illumination vanishes in correspondence with the roots of the equation j1(z) = 0. if these be called z1 z2, z3, ... the radii of the dark rings in the diffraction pattern are f[lambda]z1 f[lambda]z2 -----------, -----------, ... 2[pi]r 2[pi]r being thus _inversely_ proportional to r. the integrations may also be effected by means of polar co-ordinates, taking first the integration with respect to [phi] so as to obtain the result for an infinitely thin annular aperture. thus, if x = [rho] cos [phi], y = [rho] sin [phi], _ _ _r _2[pi] / / / / c = | | cos px dx dy = | | cos (p[rho] cos [theta]) [rho]d[rho] d[theta]. _/_/ _/0 _/0 now by definition _1⁄2[pi] 2 / z2 z^4 z^6 j0(z) = ---- | cos(z cos[theta])d[theta] = -- + ----- - -------- + ... (11). [pi] _/0 22 22·42 22·42·62 the value of c for an annular aperture of radius r and width dr is thus dc = 2 [pi]j0 (p[rho]) [rho] d[rho], (12). for the complete circle, _ pr 2[pi] / c = ----- | j0(z) zdz p2 _/0 2[pi] /p2r2 p^4 r^4 p^6 r^6 \ = ------ ( ---- - ------- + -------- - ... ) p2 \ 2 22·42 22·42·62 / 2j1(pr) = [pi]r2 · ------- as before. pr in these expressions we are to replace p by k[xi]/f, or rather, since the diffraction pattern is symmetrical, by kr/f, where r is the distance of any point in the focal plane from the centre of the system. the roots of j0(z) after the first may be found from z .050561 .053041 .262051 ---- = i - .25 + ------- - --------- + ---------- ... (13), [pi] 4i - 1 (4i - 1)3 (4i - 1)^5 and those of j1(z) from z .151982 .015399 .245835 ---- = i + .25 - ------- + --------- + ---------- ... (14), [pi] 4i + 1 (4i + 1)3 (4i + 1)^5 formulae derived by stokes (_camb. trans._, 1850, vol. ix.) from the descending series.[1] the following table gives the actual values:-- +---+--------------------+--------------------+ | | z | z | | i | ---- for j0(z) = 0 | ---- for j1(z) = 0 | | | [pi] | [pi] | +---+--------------------+--------------------+ | 1 | 7655 | 1 2197 | | 2 | 1 7571 | 2 2330 | | 3 | 2 7546 | 3 2383 | | 4 | 3 7534 | 4 2411 | | 5 | 4 7527 | 5 2428 | | 6 | 5 7522 | 6 2439 | | 7 | 6 7519 | 7 2448 | | 8 | 7 7516 | 8 2454 | | 9 | 8 7514 | 9 2459 | |10 | 9 7513 | 10 2463 | +---+--------------------+--------------------+ in both cases the image of a mathematical point is thus a symmetrical ring system. the greatest brightness is at the centre, where dc = 2[pi][rho] d[rho], c = [pi]r2. for a certain distance outwards this remains sensibly unimpaired and then gradually diminishes to zero, as the secondary waves become discrepant in phase. the subsequent revivals of brightness forming the bright rings are necessarily of inferior brilliancy as compared with the central disk. the first dark ring in the diffraction pattern of the complete circular aperture occurs when r/f = 1.2197 × [lambda]/2r (15). we may compare this with the corresponding result for a rectangular aperture of width a, [xi]/f =[lambda]/a; and it appears that in consequence of the preponderance of the central parts, the compensation in the case of the circle does not set in at so small an obliquity as when the circle is replaced by a rectangular aperture, whose side is equal to the diameter of the circle. again, if we compare the complete circle with a narrow annular aperture of the same radius, we see that in the latter case the first dark ring occurs at a much smaller obliquity, viz. r/f = .7655 × [lambda]/2r. it has been found by sir william herschel and others that the definition of a telescope is often improved by stopping off a part of the central area of the object-glass; but the advantage to be obtained in this way is in no case great, and anything like a reduction of the aperture to a narrow annulus is attended by a development of the external luminous rings sufficient to outweigh any improvement due to the diminished diameter of the central area.[2] the maximum brightnesses and the places at which they occur are easily determined with the aid of certain properties of the bessel's functions. it is known (see spherical harmonics) that j0'(z) = -j1(z), (16); 1 j2(z) = - j1(z) - j1'(z) (17); z 2 j0(z) + j2(z) = - j1(z) (18). z the maxima of c occur when d /j1(z)\ j1'(z) j1(z) -- (-------) = ------ - ----- = 0; dz \ z / z z2 or by 17 when j2(z) = 0. when z has one of the values thus determined, 2 - j1(z) = j0(z). z the accompanying table is given by lommel, in which the first column gives the roots of j2(z) = 0, and the second and third columns the corresponding values of the functions specified. if appears that the maximum brightness in the first ring is only about 1/57 of the brightness at the centre. +-------------------------------------------+ | z 2z^-1 j1(z) 4z^-2 j12(z) | +-------------------------------------------+ | | | .000000 +1.000000 1.000000 | | 5.135630 - .132279 .017498 | | 8.417236 + .064482 .004158 | | 11.619857 - .040008 .001601 | | 14.795938 + .027919 .000779 | | 17.959820 - .020905 .000437 | +-------------------------------------------+ we will now investigate the total illumination distributed over the area of the circle of radius r. we have [pi]2r^4 4j12(z) i^2 = ----------- · ------- (19), [lambda]2f2 z2 where z = 2[pi]rr/[lambda]f (20). thus _ _ _ / [lambda]2f2 / / 2[pi] | i2rdr = ----------- | i2zdz = [pi]r2·2 | z^-1 j12(z)dz. _/ 2[pi]r2 _/ _/ now by (17), (18) z^-1 j1(z) = j0(z) - j1'(z); so that d d z^-1j12(z) = 1⁄2 -- j02 - 1⁄2 -- j12(z), dz dz and _z / 2 | z^-1 j12(z)dz = 1 - j02(z) - j12(z) (21). _/0 if r, or z, be infinite, j0(z), j1(z) vanish, and the whole illumination is expressed by [pi]r2, in accordance with the general principle. in any case the proportion of the whole illumination to be found outside the circle of radius r is given by j02(z) + j12(z). for the dark rings j1(z) = 0; so that the fraction of illumination outside any dark ring is simply j02(z). thus for the first, second, third and fourth dark rings we get respectively .161, .090, .062, .047, showing that more than 9/10ths of the whole light is concentrated within the area of the second dark ring (_phil. mag._, 1881). when z is great, the descending series (10) gives 2j1(z) 2 / / 2 \ ------ = - / ( ----- ) sin(z - 1⁄4[pi]) (22); z z \/ \[pi]z/ so that the places of maxima and minima occur at equal intervals. the mean brightness varies as z^-3 (or as r^-3), and the integral found by multiplying it by zdz and integrating between 0 and [oo] converges. it may be instructive to contrast this with the case of an infinitely narrow annular aperture, where the brightness is proportional to j02(z). when z is great, / 2 j0(z) = \ / ----- cos(z^-1⁄4 [pi]). \/ [pi]z the mean brightness varies as z^-1; and the integral _ / [oo] | j02(z)z dz is not convergent. _/ 0 5. _resolving power of telescopes._--the efficiency of a telescope is of course intimately connected with the size of the disk by which it represents a mathematical point. in estimating theoretically the resolving power on a double star we have to consider the illumination of the field due to the superposition of the two independent images. if the angular interval between the components of a double star were equal to twice that expressed in equation (15) above, the central disks of the diffraction patterns would be just in contact. under these conditions there is no doubt that the star would appear to be fairly resolved, since the brightness of its external ring system is too small to produce any material confusion, unless indeed the components are of very unequal magnitude. the diminution of the star disks with increasing aperture was observed by sir william herschel, and in 1823 fraunhofer formulated the law of inverse proportionality. in investigations extending over a long series of years, the advantage of a large aperture in separating the components of close double stars was fully examined by w. r. dawes. the resolving power of telescopes was investigated also by j. b. l. foucault, who employed a scale of equal bright and dark alternate parts; it was found to be proportional to the aperture and independent of the focal length. in telescopes of the best construction and of moderate aperture the performance is not sensibly prejudiced by outstanding aberration, and the limit imposed by the finiteness of the waves of light is practically reached. m. e. verdet has compared foucault's results with theory, and has drawn the conclusion that the radius of the visible part of the image of a luminous point was equal to half the radius of the first dark ring. the application, unaccountably long delayed, of this principle to the microscope by h. l. f. helmholtz in 1871 is the foundation of the important doctrine of the _microscopic limit_. it is true that in 1823 fraunhofer, inspired by his observations upon gratings, had very nearly hit the mark.[3] and a little before helmholtz, e. abbe published a somewhat more complete investigation, also founded upon the phenomena presented by gratings. but although the argument from gratings is instructive and convenient in some respects, its use has tended to obscure the essential unity of the principle of the limit of resolution whether applied to telescopes or microscopes. [illustration: fig. 4.] in fig. 4, ab represents the axis of an optical instrument (telescope or microscope), a being a point of the object and b a point of the image. by the operation of the object-glass ll' all the rays issuing from a arrive in the same phase at b. thus if a be self-luminous, the illumination is a maximum at b, where all the secondary waves agree in phase. b is in fact the centre of the diffraction disk which constitutes the image of a. at neighbouring points the illumination is less, in consequence of the discrepancies of phase which there enter. in like manner if we take a neighbouring point p, also self-luminous, in the plane of the object, the waves which issue from it will arrive at b with phases no longer absolutely concordant, and the discrepancy of phase will increase as the interval ap increases. when the interval is very small the discrepancy, though mathematically existent, produces no practical effect; and the illumination at b due to p is as important as that due to a, the intensities of the two luminous sources being supposed equal. under these conditions it is clear that a and p are not separated in the image. the question is to what amount must the distance ap be increased in order that the difference of situation may make itself felt in the image. this is necessarily a question of degree; but it does not require detailed calculations in order to show that the discrepancy first becomes conspicuous when the phases corresponding to the various secondary waves which travel from p to b range over a complete period. the illumination at b due to p then becomes comparatively small, indeed for some forms of aperture evanescent. the extreme discrepancy is that between the waves which travel through the outermost parts of the object-glass at l and l'; so that if we adopt the above standard of resolution, the question is where must p be situated in order that the relative retardation of the rays pl and pl' may on their arrival at b amount to a wave-length ([lambda]). in virtue of the general law that the reduced optical path is stationary in value, this retardation may be calculated without allowance for the different paths pursued on the farther side of l, l', so that the value is simply pl - pl'. now since ap is very small, al' - pl' = ap sin [alpha], where [alpha] is the angular semi-aperture l'ab. in like manner pl - al has the same value, so that pl - pl' = 2ap sin [alpha]. according to the standard adopted, the condition of resolution is therefore that ap, or [epsilon], should exceed 1⁄2[lambda]/sin [alpha]. if [epsilon] be less than this, the images overlap too much; while if [epsilon] greatly exceed the above value the images become unnecessarily separated. in the above argument the whole space between the object and the lens is supposed to be occupied by matter of one refractive index, and [lambda] represents the wave-length _in this medium_ of the kind of light employed. if the restriction as to uniformity be violated, what we have ultimately to deal with is the wave-length in the medium immediately surrounding the object. calling the refractive index [mu], we have as the critical value of [epsilon], [epsilon] = 1⁄2[lambda]0/[mu] sin[alpha], (1), [lambda]0 being the wave-length _in vacuo_. the denominator [mu] sin [alpha] is the quantity well known (after abbe) as the "numerical aperture." the extreme value possible for [alpha] is a right angle, so that for the microscopic limit we have [epsilon] = 1⁄2[lambda]0/[mu] (2). the limit can be depressed only by a diminution in [lambda]0, such as photography makes possible, or by an increase in [mu], the refractive index of the medium in which the object is situated. the statement of the law of resolving power has been made in a form appropriate to the microscope, but it admits also of immediate application to the telescope. if 2r be the diameter of the object-glass and d the distance of the object, the angle subtended by ap is [epsilon]/d, and the angular resolving power is given by [lambda]/2d sin[alpha] = [lambda]/2r (3). this method of derivation (substantially due to helmholtz) makes it obvious that there is no essential difference of principle between the two cases, although the results are conveniently stated in different forms. in the case of the telescope we have to deal with a linear measure of aperture and an angular limit of resolution, whereas in the case of the microscope the limit of resolution is linear, and it is expressed in terms of angular aperture. it must be understood that the above argument distinctly assumes that the different parts of the object are self-luminous, or at least that the light proceeding from the various points is without phase relations. as has been emphasized by g. j. stoney, the restriction is often, perhaps usually, violated in the microscope. a different treatment is then necessary, and for some of the problems which arise under this head the method of abbe is convenient. the importance of the general conclusions above formulated, as imposing a limit upon our powers of direct observation, can hardly be overestimated; but there has been in some quarters a tendency to ascribe to it a more precise character than it can bear, or even to mistake its meaning altogether. a few words of further explanation may therefore be desirable. the first point to be emphasized is that nothing whatever is said as to the smallness of a single object that may be made visible. the eye, unaided or armed with a telescope, is able to see, as points of light, stars subtending no sensible angle. the visibility of a star is a question of brightness simply, and has nothing to do with resolving power. the latter element enters only when it is a question of recognizing the duplicity of a double star, or of distinguishing detail upon the surface of a planet. so in the microscope there is nothing except lack of light to hinder the visibility of an object however small. but if its dimensions be much less than the half wave-length, it can only be seen as a whole, and its parts cannot be distinctly separated, although in cases near the border line some inference may be possible, founded upon experience of what appearances are presented in various cases. interesting observations upon particles, _ultra-microscopic_ in the above sense, have been recorded by h. f. w. siedentopf and r. a. zsigmondy (_drude's ann._, 1903, 10, p. 1). in a somewhat similar way a dark linear interruption in a bright ground may be visible, although its actual width is much inferior to the half wave-length. in illustration of this fact a simple experiment may be mentioned. in front of the naked eye was held a piece of copper foil perforated by a fine needle hole. observed through this the structure of some wire gauze just disappeared at a distance from the eye equal to 17 in., the gauze containing 46 meshes to the inch. on the other hand, a single wire 0.034 in. in diameter remained fairly visible up to a distance of 20 ft. the ratio between the limiting angles subtended by the periodic structure of the gauze and the diameter of the wire was (.022/.034) × (240/17) = 9.1. for further information upon this subject reference may be made to _phil. mag._, 1896, 42, p. 167; _journ. r. micr. soc._, 1903, p. 447. 6. _coronas or glories._--the results of the theory of the diffraction patterns due to circular apertures admit of an interesting application to _coronas_, such as are often seen encircling the sun and moon. they are due to the interposition of small spherules of water, which act the part of diffracting obstacles. in order to the formation of a well-defined corona it is essential that the particles be exclusively, or preponderatingly, of one size. if the origin of light be treated as infinitely small, and be seen in focus, whether with the naked eye or with the aid of a telescope, the whole of the light in the absence of obstacles would be concentrated in the immediate neighbourhood of the focus. at other parts of the field the effect is the same, in accordance with the principle known as babinet's, whether the imaginary screen in front of the object-glass is generally transparent but studded with a number of opaque circular disks, or is generally opaque but perforated with corresponding apertures. since at these points the resultant due to the whole aperture is zero, any two portions into which the whole may be divided must give equal and opposite resultants. consider now the light diffracted in a direction many times more oblique than any with which we should be concerned, were the whole aperture uninterrupted, and take first the effect of a single small aperture. the light in the proposed direction is that determined by the size of the small aperture in accordance with the laws already investigated, and its phase depends upon the position of the aperture. if we take a direction such that the light (of given wave-length) from a single aperture vanishes, the evanescence continues even when the whole series of apertures is brought into contemplation. hence, whatever else may happen, there must be a system of dark rings formed, the same as from a single small aperture. in directions other than these it is a more delicate question how the partial effects should be compounded. if we make the extreme suppositions of an infinitely small source and absolutely homogeneous light, there is no escape from the conclusion that the light in a definite direction is arbitrary, that is, dependent upon the chance distribution of apertures. if, however, as in practice, the light be heterogeneous, the source of finite area, the obstacles in motion, and the discrimination of different directions imperfect, we are concerned merely with the mean brightness found by varying the arbitrary phase-relations, and this is obtained by simply multiplying the brightness due to a single aperture by the number of apertures (n) (see interference of light, § 4). the diffraction pattern is therefore that due to a single aperture, merely brightened n times. in his experiments upon this subject fraunhofer employed plates of glass dusted over with lycopodium, or studded with small metallic disks of uniform size; and he found that the diameters of the rings were proportional to the length of the waves and inversely as the diameter of the disks. in another respect the observations of fraunhofer appear at first sight to be in disaccord with theory; for his measures of the diameters of the red rings, visible when white light was employed, correspond with the law applicable to dark rings, and not to the different law applicable to the luminous maxima. verdet has, however, pointed out that the observation in this form is essentially different from that in which homogeneous red light is employed, and that the position of the red rings would correspond to the _absence_ of blue-green light rather than to the greatest abundance of red light. verdet's own observations, conducted with great care, fully confirm this view, and exhibit a complete agreement with theory. by measurements of coronas it is possible to infer the size of the particles to which they are due, an application of considerable interest in the case of natural coronas--the general rule being the larger the corona the smaller the water spherules. young employed this method not only to determine the diameters of cloud particles (e.g. 1/1000 in.), but also those of fibrous material, for which the theory is analogous. his instrument was called the _eriometer_ (see "chromatics," vol. iii. of supp. to _ency. brit._, 1817). 7. _influence of aberration. optical power of instruments._--our investigations and estimates of resolving power have thus far proceeded upon the supposition that there are no optical imperfections, whether of the nature of a regular aberration or dependent upon irregularities of material and workmanship. in practice there will always be a certain aberration or error of phase, which we may also regard as the deviation of the actual wave-surface from its intended position. in general, we may say that aberration is unimportant when it nowhere (or at any rate over a relatively small area only) exceeds a small fraction of the wave-length ([lamda]). thus in estimating the intensity at a focal point, where, in the absence of aberration, all the secondary waves would have exactly the same phase, we see that an aberration nowhere exceeding 1⁄4[lambda] can have but little effect. the only case in which the influence of small aberration upon the entire image has been calculated (_phil. mag._, 1879) is that of a rectangular aperture, traversed by a cylindrical wave with aberration equal to cx3. the aberration is here unsymmetrical, the wave being in advance of its proper place in one half of the aperture, but behind in the other half. no terms in x or x2 need be considered. the first would correspond to a general turning of the beam; and the second would imply imperfect focusing of the central parts. the effect of aberration may be considered in two ways. we may suppose the aperture (a) constant, and inquire into the operation of an increasing aberration; or we may take a given value of c (i.e. a given wave-surface) and examine the effect of a varying aperture. the results in the second case show that an increase of aperture up to that corresponding to an extreme aberration of half a period has no ill effect upon the central band (§ 3), but it increases unduly the intensity of one of the neighbouring lateral bands; and the practical conclusion is that the best results will be obtained from an aperture giving an extreme aberration of from a quarter to half a period, and that with an increased aperture aberration is not so much a direct cause of deterioration as an obstacle to the attainment of that improved definition which should accompany the increase of aperture. if, on the other hand, we suppose the aperture given, we find that aberration begins to be distinctly mischievous when it amounts to about a quarter period, i.e. when the wave-surface deviates at each end by a quarter wave-length from the true plane. as an application of this result, let us investigate what amount of temperature disturbance in the tube of a telescope may be expected to impair definition. according to j. b. biot and f. j. d. arago, the index [mu] for air at t° c. and at atmospheric pressure is given by .00029 [mu] - 1 = -----------. 1 + .0037 t if we take 0° c. as standard temperature, [delta][mu] = -1.1 × 10^-6. thus, on the supposition that the irregularity of temperature t extends through a length l, and produces an acceleration of a quarter of a wave-length, 1⁄4[lambda] = 1.1 lt × 10^-6; or, if we take [lambda] = 5.3 × 10^-5, lt = 12, the unit of length being the centimetre. we may infer that, in the case of a telescope tube 12 cm. long, a stratum of air heated 1° c. lying along the top of the tube, and occupying a moderate fraction of the whole volume, would produce a not insensible effect. if the change of temperature progressed uniformly from one side to the other, the result would be a lateral displacement of the image without loss of definition; but in general both effects would be observable. in longer tubes a similar disturbance would be caused by a proportionally less difference of temperature. s. p. langley has proposed to obviate such ill-effects by stirring the air included within a telescope tube. it has long been known that the definition of a carbon bisulphide prism may be much improved by a vigorous shaking. we will now consider the application of the principle to the formation of images, unassisted by reflection or refraction (_phil. mag._, 1881). the function of a lens in forming an image is to compensate by its variable thickness the differences of phase which would otherwise exist between secondary waves arriving at the focal point from various parts of the aperture. if we suppose the diameter of the lens to be given (2r), and its focal length f gradually to increase, the original differences of phase at the image of an infinitely distant luminous point diminish without limit. when f attains a certain value, say f1, the extreme error of phase to be compensated falls to 1⁄4[lambda]. but, as we have seen, such an error of phase causes no sensible deterioration in the definition; so that from this point onwards the lens is useless, as only improving an image already sensibly as perfect as the aperture admits of. throughout the operation of increasing the focal length, the resolving power of the instrument, which depends only upon the aperture, remains unchanged; and we thus arrive at the rather startling conclusion that a telescope of any degree of resolving power might be constructed without an object-glass, if only there were no limit to the admissible focal length. this last proviso, however, as we shall see, takes away almost all practical importance from the proposition. to get an idea of the magnitudes of the quantities involved, let us take the case of an aperture of 1/5 in., about that of the pupil of the eye. the distance f1, which the actual focal length must exceed, is given by / \/ (f12 + r2) - f1 = 1⁄4[lambda]; so that f1 = 2r2/[lambda] (1). thus, if [lambda] = 1/4000, r = 1/10, we find f1 = 800 inches. the image of the sun thrown upon a screen at a distance exceeding 66 ft., through a hole 1/5 in. in diameter, is therefore at least as well defined as that seen direct. as the minimum focal length increases with the square of the aperture, a quite impracticable distance would be required to rival the resolving power of a modern telescope. even for an aperture of 4 in., f1 would have to be 5 miles. a similar argument may be applied to find at what point an achromatic lens becomes sensibly superior to a single one. the question is whether, when the adjustment of focus is correct for the central rays of the spectrum, the error of phase for the most extreme rays (which it is necessary to consider) amounts to a quarter of a wave-length. if not, the substitution of an achromatic lens will be of no advantage. calculation shows that, if the aperture be 1/5 in., an achromatic lens has no sensible advantage if the focal length be greater than about 11 in. if we suppose the focal length to be 66 ft., a single lens is practically perfect up to an aperture of 1.7 in. another obvious inference from the necessary imperfection of optical images is the uselessness of attempting anything like an absolute destruction of spherical aberration. an admissible error of phase of 1⁄4[lambda] will correspond to an error of 1/8[lambda] in a reflecting and 1⁄2[lambda] in a (glass) refracting surface, the incidence in both cases being perpendicular. if we inquire what is the greatest admissible longitudinal aberration ([delta]f) in an object-glass according to the above rule, we find [delta]f = [lambda][alpha]^-2 (2), [alpha] being the angular semi-aperture. in the case of a single lens of glass with the most favourable curvatures, [delta]f is about equal to [alpha]2f, so that [alpha]^4 must not exceed [lambda]/f. for a lens of 3 ft. focus this condition is satisfied if the aperture does not exceed 2 in. when parallel rays fall directly upon a spherical mirror the longitudinal aberration is only about one-eighth as great as for the most favourably shaped single lens of equal focal length and aperture. hence a spherical mirror of 3 ft. focus might have an aperture of 21⁄2 in., and the image would not suffer materially from aberration. on the same principle we may estimate the least visible displacement of the eye-piece of a telescope focused upon a distant object, a question of interest in connexion with range-finders. it appears (_phil. mag._, 1885, 20, p. 354) that a displacement [delta]f from the true focus will not sensibly impair definition, provided [delta]f < f2[lambda]/r2 (3), 2r being the diameter of aperture. the linear accuracy required is thus a function of the _ratio_ of aperture to focal length. the formula agrees well with experiment. the principle gives an instantaneous solution of the question of the ultimate optical efficiency in the method of "mirror-reading," as commonly practised in various physical observations. a rotation by which one edge of the mirror advances 1⁄4[lambda] (while the other edge retreats to a like amount) introduces a phase-discrepancy of a whole period where before the rotation there was complete agreement. a rotation of this amount should therefore be easily visible, but the limits of resolving power are being approached; and the conclusion is independent of the focal length of the mirror, and of the employment of a telescope, provided of course that the reflected image is seen in focus, and that the full width of the mirror is utilized. a comparison with the method of a material pointer, attached to the parts whose rotation is under observation, and viewed through a microscope, is of interest. the limiting efficiency of the microscope is attained when the angular aperture amounts to 180°; and it is evident that a lateral displacement of the point under observation through 1⁄2[lambda] entails (at the old image) a phase-discrepancy of a whole period, one extreme ray being accelerated and the other retarded by half that amount. we may infer that the limits of efficiency in the two methods are the same when the length of the pointer is equal to the width of the mirror. [illustraton: fig. 5.] we have seen that in perpendicular reflection a surface error not exceeding 1/8[lambda] may be admissible. in the case of oblique reflection at an angle [phi], the error of retardation due to an elevation bd (fig. 5) is qq' - qs = bd sec [phi](1 - cos sqq') = bd sec [phi] (1 + cos 2[phi]) = 2bd cos [phi]; from which it follows that an error of given magnitude in the figure of a surface is less important in oblique than in perpendicular reflection. it must, however, be borne in mind that errors can sometimes be compensated by altering adjustments. if a surface intended to be flat is affected with a slight general curvature, a remedy may be found in an alteration of focus, and the remedy is the less complete as the reflection is more oblique. the formula expressing the optical power of prismatic spectroscopes may readily be investigated upon the principles of the wave theory. let a0b0 be a plane wave-surface of the light before it falls upon the prisms, ab the corresponding wave-surface for a particular part of the spectrum after the light has passed the prisms, or after it has passed the eye-piece of the observing telescope. the path of a ray from the wave-surface a0b0 to a or b is determined by the condition that the optical distance, [int] [mu]ds, is a minimum; and, as ab is by supposition a wave-surface, this optical distance is the same for both points. thus _ _ / / | [mu]ds (for a) = | [mu]ds (for b) (4). _/ _/ we have now to consider the behaviour of light belonging to a neighbouring part of the spectrum. the path of a ray from the wave-surface a0b0 to the point a is changed; but in virtue of the minimum property the change may be neglected in calculating the optical distance, as it influences the result by quantities of the second order only in the changes of refrangibility. accordingly, the optical distance from a0b0 to a is represented by [int]([mu] + [delta][mu])ds, the integration being along the original path a0 ... a; and similarly the optical distance between a0b0 and b is represented by [int] ([mu] + [delta][mu])ds, the integration being along b0 ... b. in virtue of (4) the difference of the optical distances to a and b is _ _ / / | [delta][mu]ds (along b0 ... b) - | [delta][mu]ds (along a0 ... a) (5). _/ _/ the new wave-surface is formed in such a position that the optical distance is constant; and therefore the _dispersion_, or the angle through which the wave-surface is turned by the change of refrangibility, is found simply by dividing (5) by the distance ab. if, as in common flint-glass spectroscopes, there is only one dispersing substance, [int] [delta][mu] ds = [delta][mu]·s, where s is simply the thickness traversed by the ray. if t2 and t1 be the thicknesses traversed by the extreme rays, and a denote the width of the emergent beam, the dispersion [theta] is given by [theta] = [delta][mu](t2 - t1)/a, or, if t1 be negligible, [theta] = [delta][mu]t/a (6). the condition of resolution of a double line whose components subtend an angle [theta] is that [theta] must exceed [lambda]/a. hence, in order that a double line may be resolved whose components have indices [mu] and [mu] + [delta][mu], it is necessary that t should exceed the value given by the following equation:-- t = [lambda]/[delta][mu] (7). 8. _diffraction gratings._--under the heading "colours of striated surfaces," thomas young (_phil. trans._, 1802) in his usual summary fashion gave a general explanation of these colours, including the law of sines, the striations being supposed to be straight, parallel and equidistant. later, in his article "chromatics" in the supplement to the 5th edition of this encyclopaedia, he shows that the colours "lose the mixed character of periodical colours, and resemble much more the ordinary prismatic spectrum, with intervals completely dark interposed," and explains it by the consideration that any phase-difference which may arise at neighbouring striae is multiplied in proportion to the total number of striae. the theory was further developed by a. j. fresnel (1815), who gave a formula equivalent to (5) below. but it is to j. von fraunhofer that we owe most of our knowledge upon this subject. his recent discovery of the "fixed lines" allowed a precision of observation previously impossible. he constructed gratings up to 340 periods to the inch by straining fine wire over screws. subsequently he ruled gratings on a layer of gold-leaf attached to glass, or on a layer of grease similarly supported, and again by attacking the glass itself with a diamond point. the best gratings were obtained by the last method, but a suitable diamond point was hard to find, and to preserve. observing through a telescope with light perpendicularly incident, he showed that the position of any ray was dependent only upon the grating interval, viz. the distance from the centre of one wire or line to the centre of the next, and not otherwise upon the thickness of the wire and the magnitude of the interspace. in different gratings the lengths of the spectra and their distances from the axis were inversely proportional to the grating interval, while with a given grating the distances of the various spectra from the axis were as 1, 2, 3, &c. to fraunhofer we owe the first accurate measurements of wave-lengths, and the method of separating the overlapping spectra by a prism dispersing in the perpendicular direction. he described also the complicated patterns seen when a point of light is viewed through two superposed gratings, whose lines cross one another perpendicularly or obliquely. the above observations relate to transmitted light, but fraunhofer extended his inquiry to the light _reflected_. to eliminate the light returned from the hinder surface of an engraved grating, he covered it with a black varnish. it then appeared that under certain angles of incidence parts of the resulting spectra were _completely polarized_. these remarkable researches of fraunhofer, carried out in the years 1817-1823, are republished in his _collected writings_ (munich, 1888). the principle underlying the action of gratings is identical with that discussed in § 2, and exemplified in j. l. soret's "zone plates." the alternate fresnel's zones are blocked out or otherwise modified; in this way the original compensation is upset and a revival of light occurs in unusual directions. if the source be a point or a line, and a collimating lens be used, the incident waves may be regarded as plane. if, further, on leaving the grating the light be received by a focusing lens, e.g. the object-glass of a telescope, the fresnel's zones are reduced to parallel and equidistant straight strips, which at certain angles coincide with the ruling. the directions of the lateral spectra are such that the passage from one element of the grating to the corresponding point of the next implies a retardation of an integral number of wave-lengths. if the grating be composed of alternate transparent and opaque parts, the question may be treated by means of the general integrals (§ 3) by merely limiting the integration to the transparent parts of the aperture. for an investigation upon these lines the reader is referred to airy's _tracts_, to verdet's _lecons_, or to r. w. wood's _physical optics_. if, however, we assume the theory of a simple rectangular aperture (§ 3); the results of the ruling can be inferred by elementary methods, which are perhaps more instructive. apart from the ruling, we know that the image of a mathematical line will be a series of narrow bands, of which the central one is by far the brightest. at the middle of this band there is complete agreement of phase among the secondary waves. the dark lines which separate the bands are the places at which the phases of the secondary wave range over an integral number of periods. if now we suppose the aperture ab to be covered by a great number of opaque strips or bars of width d, separated by transparent intervals of width a, the condition of things in the directions just spoken of is not materially changed. at the central point there is still complete agreement of phase; but the amplitude is diminished in the ratio of a : a + d. in another direction, making a small angle with the last, such that the projection of ab upon it amounts to a few wave-lengths, it is easy to see that the mode of interference is the same as if there were no ruling. for example, when the direction is such that the projection of ab upon it amounts to one wave-length, the elementary components neutralize one another, because their phases are distributed symmetrically, though discontinuously, round the entire period. the only effect of the ruling is to diminish the amplitude in the ratio a : a + d; and, except for the difference in illumination, the appearance of a line of light is the same as if the aperture were perfectly free. the lateral (spectral) images occur in such directions that the projection of the element (a + d) of the grating upon them is an exact multiple of [lambda]. the effect of each of the n elements of the grating is then the same; and, unless this vanishes on account of a particular adjustment of the ratio a : d, the resultant amplitude becomes comparatively very great. these directions, in which the retardation between a and b is exactly mn[lambda], may be called the principal directions. on either side of any one of them the illumination is distributed according to the same law as for the central image (m = 0), vanishing, for example, when the retardation amounts to (mn ± 1)[lambda]. in considering the relative brightnesses of the different spectra, it is therefore sufficient to attend merely to the principal directions, provided that the whole deviation be not so great that its cosine differs considerably from unity. we have now to consider the amplitude due to a single element, which we may conveniently regard as composed of a transparent part a bounded by two opaque parts of width 1⁄2d. the phase of the resultant effect is by symmetry that of the component which comes from the middle of a. the fact that the other components have phases differing from this by amounts ranging between ± am[pi]/(a + d) causes the resultant amplitude to be less than for the central image (where there is complete phase agreement). if bm denote the brightness of the m^th lateral image, and b0 that of the central image, we have _ _+ am[pi]/(a + d) _ | / 2am[pi] |2 /a + d \2 am[pi] b_m : b0 = | | cosx dx ÷ ------- | = ( ------ ) sin2 ------ (1). |_ _/ a + d _| \am[pi]/ a + d -am[pi]/(a + d) if b denotes the brightness of the central image when the whole of the space occupied by the grating is transparent, we have b0 : b = a2 : (a + d)2, and thus 1 am[pi] bm : b = ------- sin2 ------ (2). m2[pi]2 a + d the sine of an angle can never be greater than unity; and consequently under the most favourable circumstances only 1/m2[pi]2 of the original light can be obtained in the m^th spectrum. we conclude that, with a grating composed of transparent and opaque parts, the utmost light obtainable in any one spectrum is in the first, and there amounts to 1/[pi]2, or about 1/10, and that for this purpose a and d must be equal. when d = a the general formula becomes sin2 1⁄2m[pi] bm : b = ----------- (3), m2[pi]2 showing that, when m is even, bm vanishes, and that, when m is odd, bm : b = 1/m2[pi]2. the third spectrum has thus only 1/9 of the brilliancy of the first. another particular case of interest is obtained by supposing a small relatively to (a + d). unless the spectrum be of very high order, we have simply bm : b = a/(a + d)2 (4); so that the brightnesses of all the spectra are the same. the light stopped by the opaque parts of the grating, together with that distributed in the central image and lateral spectra, ought to make up the brightness that would be found in the central image, were all the apertures transparent. thus, if a = d, we should have 1 1 2 / 1 1 \ 1 = - + - + ----- ( 1 + - + -- + ... ), 2 4 [pi]2 \ 9 25 / which is true by a known theorem. in the general case ___m=[oo] a / a \2 2 \ 1 /m[pi]a\ ----- = ( ----- ) + ----- > -- sin2( ------ ), a + d \a + d/ [pi]2 /__ m2 \ a + d/ m=1 a formula which may be verified by fourier's theorem. according to a general principle formulated by j. babinet, the brightness of a lateral spectrum is not affected by an interchange of the transparent and opaque parts of the grating. the vibrations corresponding to the two parts are precisely antagonistic, since if both were operative the resultant would be zero. so far as the application to gratings is concerned, the same conclusion may be derived from (2). [illustration: fig. 6.] from the value of bm : b0 we see that no lateral spectrum can surpass the central image in brightness; but this result depends upon the hypothesis that the ruling acts by opacity, which is generally very far from being the case in practice. in an engraved glass grating there is no opaque material present by which light could be absorbed, and the effect depends upon a difference of retardation in passing the alternate parts. it is possible to prepare gratings which give a lateral spectrum brighter than the central image, and the explanation is easy. for if the alternate parts were equal and alike transparent, but so constituted as to give a relative retardation of 1⁄2[lambda], it is evident that the central image would be entirely extinguished, while the first spectrum would be four times as bright as if the alternate parts were opaque. if it were possible to introduce at every part of the aperture of the grating an arbitrary retardation, all the light might be concentrated in any desired spectrum. by supposing the retardation to vary uniformly and continuously we fall upon the case of an ordinary prism: but there is then no diffraction spectrum in the usual sense. to obtain such it would be necessary that the retardation should gradually alter by a wave-length in passing over any element of the grating, and then fall back to its previous value, thus springing suddenly over a wave-length (_phil. mag._, 1874, 47, p. 193). it is not likely that such a result will ever be fully attained in practice; but the case is worth stating, in order to show that there is no theoretical limit to the concentration of light of assigned wave-length in one spectrum, and as illustrating the frequently observed unsymmetrical character of the spectra on the two sides of the central image.[4] we have hitherto supposed that the light is incident perpendicularly upon the grating; but the theory is easily extended. if the incident rays make an angle [theta] with the normal (fig. 6), and the diffracted rays make an angle [phi] (upon the same side), the relative retardation from each element of width (a + d) to the next is (a + d) (sin[theta] + sin[phi]); and this is the quantity which is to be equated to m[lambda]. thus sin[theta] + sin[phi] = 2 sin 1⁄2([theta] + [phi]) cos 1⁄2([theta] - [phi]) = m[lambda]/(a + d) (5). the "deviation" is ([theta] + [phi]), and is therefore a minimum when [theta] = [phi], i.e. when the grating is so situated that the angles of incidence and diffraction are equal. in the case of a reflection grating the same method applies. if [theta] and [phi] denote the angles with the normal made by the incident and diffracted rays, the formula (5) still holds, and, if the deviation be reckoned from the direction of the regularly reflected rays, it is expressed as before by ([theta] + [phi]), and is a minimum when [theta] = [phi], that is, when the diffracted rays return upon the course of the incident rays. [illustration: fig. 7.] in either case (as also with a prism) the position of minimum deviation leaves the width of the beam unaltered, i.e. neither magnifies nor diminishes the angular width of the object under view. from (5) we see that, when the light falls perpendicularly upon a grating ([theta] = 0), there is no spectrum formed (the image corresponding to m = 0 not being counted as a spectrum), if the grating interval [sigma] or (a + d) is less than [lambda]. under these circumstances, if the material of the grating be completely transparent, the whole of the light must appear in the direct image, and the ruling is not perceptible. from the absence of spectra fraunhofer argued that there must be a microscopic limit represented by [lambda]; and the inference is plausible, to say the least (_phil. mag._, 1886). fraunhofer should, however, have fixed the microscopic limit at 1⁄2[lambda], as appears from (5), when we suppose [theta] = 1⁄2[pi], [phi] = 1⁄2[pi]. [illustration: fig. 8.] we will now consider the important subject of the resolving power of gratings, as dependent upon the number of lines (n) and the order of the spectrum observed (m). let bp (fig. 8) be the direction of the principal maximum (middle of central band) for the wave-length [lambda] in the m^th spectrum. then the relative retardation of the extreme rays (corresponding to the edges a, b of the grating) is mn[lambda]. if bq be the direction for the first minimum (the darkness between the central and first lateral band), the relative retardation of the extreme rays is (mn + 1)[lambda]. suppose now that [lambda] + [delta][lambda] is the wave-length for which bq gives the principal maximum, then (mn + 1)[lambda] = mn([lambda] + [delta][lambda]); whence [delta][lambda]/[lambda] = 1/mn (6). according to our former standard, this gives the smallest difference of wave-lengths in a double line which can be just resolved; and we conclude that the resolving power of a grating depends only upon the total number of lines, and upon the order of the spectrum, without regard to any other considerations. it is here of course assumed that the n lines are really utilized. in the case of the d lines the value of [delta][lambda]/[lambda] is about 1/1000; so that to resolve this double line in the first spectrum requires 1000 lines, in the second spectrum 500, and so on. it is especially to be noticed that the resolving power does not depend directly upon the closeness of the ruling. let us take the case of a grating 1 in. broad, and containing 1000 lines, and consider the effect of interpolating an additional 1000 lines, so as to bisect the former intervals. there will be destruction by interference of the first, third and odd spectra generally; while the advantage gained in the spectra of even order is not in dispersion, nor in resolving power, but simply in brilliancy, which is increased four times. if we now suppose half the grating cut away, so as to leave 1000 lines in half an inch, the dispersion will not be altered, while the brightness and resolving power are halved. there is clearly no theoretical limit to the resolving power of gratings, even in spectra of given order. but it is possible that, as suggested by rowland,[5] the structure of natural spectra may be too coarse to give opportunity for resolving powers much higher than those now in use. however this may be, it would always be possible, with the aid of a grating of given resolving power, to construct artificially from white light mixtures of slightly different wave-length whose resolution or otherwise would discriminate between powers inferior and superior to the given one.[6] if we define as the "dispersion" in a particular part of the spectrum the ratio of the angular interval d[theta] to the corresponding increment of wave-length d[lambda], we may express it by a very simple formula. for the alteration of wave-length entails, at the two limits of a diffracted wave-front, a relative retardation equal to mnd[lambda]. hence, if a be the width of the diffracted beam, and d[theta] the angle through which the wave-front is turned, ad[theta] = mn d[lambda], or dispersion = mn/a (7). the resolving power and the width of the emergent beam fix the optical character of the instrument. the latter element must eventually be decreased until less than the diameter of the pupil of the eye. hence a wide beam demands treatment with further apparatus (usually a telescope) of high magnifying power. in the above discussion it has been supposed that the ruling is accurate, and we have seen that by increase of m a high resolving power is attainable with a moderate number of lines. but this procedure (apart from the question of illumination) is open to the objection that it makes excessive demands upon accuracy. according to the principle already laid down it can make but little difference in the principal direction corresponding to the first spectrum, provided each line lie within a quarter of an interval (a + d) from its theoretical position. but, to obtain an equally good result in the m^th spectrum, the error must be less than 1/m of the above amount.[7] there are certain errors of a systematic character which demand special consideration. the spacing is usually effected by means of a screw, to each revolution of which corresponds a large number (e.g. one hundred) of lines. in this way it may happen that although there is almost perfect periodicity with each revolution of the screw after (say) 100 lines, yet the 100 lines themselves are not equally spaced. the "ghosts" thus arising were first described by g. h. quincke (_pogg. ann._, 1872, 146, p. 1), and have been elaborately investigated by c. s. peirce (_ann. journ. math._, 1879, 2, p. 330), both theoretically and experimentally. the general nature of the effects to be expected in such a case may be made clear by means of an illustration already employed for another purpose. suppose two similar and accurately ruled transparent gratings to be superposed in such a manner that the lines are parallel. if the one set of lines exactly bisect the intervals between the others, the grating interval is practically halved, and the previously existing spectra of odd order vanish. but a very slight relative displacement will cause the apparition of the odd spectra. in this case there is approximate periodicity in the half interval, but complete periodicity only after the whole interval. the advantage of approximate bisection lies in the superior brilliancy of the surviving spectra; but in any case the compound grating may be considered to be perfect in the longer interval, and the definition is as good as if the bisection were accurate. [illustration: | | | | | ( ( ( | | | | ) | ( fig. 9.--x2. fig. 10.--y2. fig. 11.--x3. fig. 12.--xy2. / / / \ | | / | \ | | | | / / / fig. 13.--xy. fig. 14.--x2y. fig. 15.--y3.] the effect of a gradual increase in the interval (fig. 9) as we pass across the grating has been investigated by m. a. cornu (_c.r._, 1875, 80, p. 655), who thus explains an anomaly observed by e. e. n. mascart. the latter found that certain gratings exercised a converging power upon the spectra formed upon one side, and a corresponding diverging power upon the spectra on the other side. let us suppose that the light is incident perpendicularly, and that the grating interval increases from the centre towards that edge which lies nearest to the spectrum under observation, and decreases towards the hinder edge. it is evident that the waves from _both_ halves of the grating are accelerated in an increasing degree, as we pass from the centre outwards, as compared with the phase they would possess were the central value of the grating interval maintained throughout. the irregularity of spacing has thus the effect of a convex lens, which accelerates the marginal relatively to the central rays. on the other side the effect is reversed. this kind of irregularity may clearly be present in a degree surpassing the usual limits, without loss of definition, when the telescope is focused so as to secure the best effect. it may be worth while to examine further the other variations from correct ruling which correspond to the various terms expressing the deviation of the wave-surface from a perfect plane. if x and y be co-ordinates in the plane of the wave-surface, the axis of y being parallel to the lines of the grating, and the origin corresponding to the centre of the beam, we may take as an approximate equation to the wave-surface x2 y2 z = ------ + bxy + ------- + [alpha]x3 + [beta]x2y + [gamma]xy2 + [delta]y3 + ... (8); 2[rho] 2[rho]' and, as we have just seen, the term in x2 corresponds to a linear error in the spacing. in like manner, the term in y2 corresponds to a general _curvature_ of the lines (fig. 10), and does not influence the definition at the (primary) focus, although it may introduce astigmatism.[8] if we suppose that everything is symmetrical on the two sides of the primary plane y = 0, the coefficients b, [beta], [delta] vanish. in spite of any inequality between [rho] and [rho]', the definition will be good to this order of approximation, provided [alpha] and [gamma] vanish. the former measures the _thickness_ of the primary focal line, and the latter measures its _curvature_. the error of ruling giving rise to [alpha] is one in which the intervals increase or decrease in _both_ directions from the centre outwards (fig. 11), and it may often be compensated by a slight rotation in azimuth of the object-glass of the observing telescope. the term in [gamma] corresponds to a _variation_ of curvature in crossing the grating (fig. 12). when the plane zx is not a plane of symmetry, we have to consider the terms in xy, x2y, and y3. the first of these corresponds to a deviation from parallelism, causing the interval to alter gradually as we pass _along_ the lines (fig. 13). the error thus arising may be compensated by a rotation of the object-glass about one of the diameters y = ± x. the term in x2y corresponds to a deviation from parallelism in the same direction on both sides of the central line (fig. 14); and that in y3 would be caused by a curvature such that there is a point of inflection at the middle of each line (fig. 15). all the errors, except that depending on [alpha], and especially those depending on [gamma] and [delta], can be diminished, without loss of resolving power, by contracting the _vertical_ aperture. a linear error in the spacing, and a general curvature of the lines, are eliminated in the ordinary use of a grating. the explanation of the difference of focus upon the two sides as due to unequal spacing was verified by cornu upon gratings purposely constructed with an increasing interval. he has also shown how to rule a plane surface with lines so disposed that the grating shall of itself give well-focused spectra. [illustration: fig. 16.] a similar idea appears to have guided h. a. rowland to his brilliant invention of concave gratings, by which spectra can be photographed without any further optical appliance. in these instruments the lines are ruled upon a spherical surface of speculum metal, and mark the intersections of the surface by a system of parallel and equidistant planes, of which the middle member passes through the centre of the sphere. if we consider for the present only the primary plane of symmetry, the figure is reduced to two dimensions. let ap (fig. 16) represent the surface of the grating, o being the centre of the circle. then, if q be any radiant point and q' its image (primary focus) in the spherical mirror ap, we have 1 1 2cos[phi] -- + - = ---------, v1 u a where v1 = aq', u = aq, a = oa, [phi] = angle of incidence qao, equal to the angle of reflection q'ao. if q be on the circle described upon oa as diameter, so that u = a cos [phi], then q' lies also upon the same circle; and in this case it follows from the symmetry that the unsymmetrical aberration (depending upon a) vanishes. this disposition is adopted in rowland's instrument; only, in addition to the central image formed at the angle [phi]' = [phi], there are a series of spectra with various values of [phi]', but all disposed upon the same circle. rowland's investigation is contained in the paper already referred to; but the following account of the theory is in the form adopted by r. t. glazebrook (_phil. mag._, 1883). in order to find the difference of optical distances between the courses qaq', qpq', we have to express qp - qa, pq' - aq'. to find the former, we have, if oaq = [phi], aop = [omega], qp2 = u2 + 4a2sin21⁄2[omega] - 4au sin 1⁄2[omega] sin (1⁄2[omega] - [phi]) = (u + a sin[phi] sin[omega])2 - a2 sin2[phi] sin2[omega] + 4a sin2 1⁄2[omega](a - u cos[phi]). now as far as [omega]^4 4 sin2 1⁄2[omega] = sin2[omega] + 1⁄4sin^4[omega], and thus to the same order qp2 = (u + a sin [phi] sin [omega])2 -a cos [phi](u - a cos [phi]) sin2[omega] + 1⁄4 a(a - u cos[phi]) sin^4 [omega]. but if we now suppose that q lies on the circle u = a cos [phi], the middle term vanishes, and we get, correct as far as [omega]^4, / / a2 sin2[phi] sin^4[omega]\ qp = (u + a sin[phi] sin[omega]) / ( 1 + ------------------------- ); \/ \ 4u / so that qp - u = a sin [phi] sin [omega] + 1/8 a sin[phi] tan[phi] sin^4 [omega] (9), in which it is to be noticed that the adjustment necessary to secure the disappearance of sin2[omega] is sufficient also to destroy the term in sin3[omega]. a similar expression can be found for q'p - q'a; and thus, if q'a = v, q'ao = [phi]', where v = a cos [phi]', we get qp + pq' - qa -aq' = a sin[omega] (sin[phi] - sin[phi]') + 1/8 a sin^4 [omega] (sin[phi] tan[phi] + sin[phi]' tan[phi]') (10). if [phi]' = [phi], the term of the first order vanishes, and the reduction of the difference of path _via_ p and _via_ a to a term of the fourth order proves not only that q and q' are conjugate foci, but also that the foci are exempt from the most important term in the aberration. in the present application [phi]' is not necessarily equal to [phi]; but if p correspond to a line upon the grating, the difference of retardations for consecutive positions of p, so far as expressed by the term of the first order, will be equal to [-+] m[lambda] (m integral), and therefore without influence, provided [sigma] (sin[phi] - sin[phi]') = ± m[lambda] (11), where [sigma] denotes the constant interval between the planes containing the lines. this is the ordinary formula for a reflecting plane grating, and it shows that the spectra are formed in the usual directions. they are here focused (so far as the rays in the primary plane are concerned) upon the circle oq'a, and the outstanding aberration is of the fourth order. in order that a large part of the field of view may be in focus at once, it is desirable that the locus of the focused spectrum should be nearly perpendicular to the line of vision. for this purpose rowland places the eye-piece at o, so that [phi] = 0, and then by (11) the value of [phi]' in the m^th spectrum is [sigma] sin [phi]' = ± m[lambda] (12). if [omega] now relate to the edge of the grating, on which there are altogether n lines, n[sigma] = 2a sin [omega], and the value of the last term in (10) becomes 1/16 n[sigma] sin3[omega] sin[phi]' tan[phi]', or 1/16 mn[lambda] sin3[omega] tan [phi]' (13). this expresses the retardation of the extreme relatively to the central ray, and is to be reckoned positive, whatever may be the signs of [omega], and [phi]'. if the semi-angular aperture ([omega]) be 1/100, and tan [phi]' = 1, mn might be as great as four millions before the error of phase would reach 1⁄4[lambda]. if it were desired to use an angular aperture so large that the aberration according to (13) would be injurious, rowland points out that on his machine there would be no difficulty in applying a remedy by making [sigma] slightly variable towards the edges. or, retaining [sigma] constant, we might attain compensation by so polishing the surface as to bring the circumference slightly forward in comparison with the position it would occupy upon a true sphere. it may be remarked that these calculations apply to the rays in the primary plane only. the image is greatly affected with astigmatism; but this is of little consequence, if [gamma] in (8) be small enough. curvature of the primary focal line having a very injurious effect upon definition, it may be inferred from the excellent performance of these gratings that [gamma] is in fact small. its value does not appear to have been calculated. the other coefficients in (8) vanish in virtue of the symmetry. the mechanical arrangements for maintaining the focus are of great simplicity. the grating at a and the eye-piece at o are rigidly attached to a bar ao, whose ends rest on carriages, moving on rails oq, aq at right angles to each other. a tie between the middle point of the rod oa and q can be used if thought desirable. the absence of chromatic aberration gives a great advantage in the comparison of overlapping spectra, which rowland has turned to excellent account in his determinations of the relative wave-lengths of lines in the solar spectrum (_phil. mag._, 1887). for absolute determinations of wave-lengths plane gratings are used. it is found (bell, _phil. mag._, 1887) that the angular measurements present less difficulty than the comparison of the grating interval with the standard metre. there is also some uncertainty as to the actual temperature of the grating when in use. in order to minimize the heating action of the light, it might be submitted to a preliminary prismatic analysis before it reaches the slit of the spectrometer, after the manner of helmholtz. in spite of the many improvements introduced by rowland and of the care with which his observations were made, recent workers have come to the conclusion that errors of unexpected amount have crept into his measurements of wave-lengths, and there is even a disposition to discard the grating altogether for fundamental work in favour of the so-called "interference methods," as developed by a. a. michelson, and by c. fabry and j. b. perot. the grating would in any case retain its utility for the reference of new lines to standards otherwise fixed. for such standards a relative accuracy of at least one part in a million seems now to be attainable. since the time of fraunhofer many skilled mechanicians have given their attention to the ruling of gratings. those of nobert were employed by a. j. angstrom in his celebrated researches upon wave-lengths. l. m. rutherfurd introduced into common use the reflection grating, finding that speculum metal was less trying than glass to the diamond point, upon the permanence of which so much depends. in rowland's dividing engine the screws were prepared by a special process devised by him, and the resulting gratings, plane and concave, have supplied the means for much of the best modern optical work. it would seem, however, that further improvements are not excluded. there are various copying processes by which it is possible to reproduce an original ruling in more or less perfection. the earliest is that of quincke, who coated a glass grating with a chemical silver deposit, subsequently thickened with copper in an electrolytic bath. the metallic plate thus produced formed, when stripped from its support, a reflection grating reproducing many of the characteristics of the original. it is best to commence the electrolytic thickening in a silver acetate bath. at the present time excellent reproductions of rowland's speculum gratings are on the market (thorp, ives, wallace), prepared, after a suggestion of sir david brewster, by coating the original with a varnish, e.g. of celluloid. much skill is required to secure that the film when stripped shall remain undeformed. a much easier method, applicable to glass originals, is that of photographic reproduction by contact printing. in several papers dating from 1872, lord rayleigh (see _collected papers_, i. 157, 160, 199, 504; iv. 226) has shown that success may be attained by a variety of processes, including bichromated gelatin and the old bitumen process, and has investigated the effect of imperfect approximation during the exposure between the prepared plate and the original. for many purposes the copies, containing lines up to 10,000 to the inch, are not inferior. it is to be desired that transparent gratings should be obtained from first-class ruling machines. to save the diamond point it might be possible to use something softer than ordinary glass as the material of the plate. 9. _talbot's bands._--these very remarkable bands are seen under certain conditions when a tolerably pure spectrum is regarded with the naked eye, or with a telescope, _half the aperture being covered by a thin plate_, e.g. _of glass or mica_. the view of the matter taken by the discoverer (_phil. mag._, 1837, 10, p. 364) was that any ray which suffered in traversing the plate a retardation of an odd number of half wave-lengths would be extinguished, and that thus the spectrum would be seen interrupted by a number of dark bars. but this explanation cannot be accepted as it stands, being open to the same objection as arago's theory of stellar scintillation.[9] it is as far as possible from being true that a body emitting homogeneous light would disappear on merely covering half the aperture of vision with a half-wave plate. such a conclusion would be in the face of the principle of energy, which teaches plainly that the retardation in question leaves the aggregate brightness unaltered. the actual formation of the bands comes about in a very curious way, as is shown by a circumstance first observed by brewster. when the retarding plate is held on the side towards the red of the spectrum, _the bands are not seen_. even in the contrary case, the thickness of the plate must not exceed a certain limit, dependent upon the purity of the spectrum. a satisfactory explanation of these bands was first given by airy (_phil. trans._, 1840, 225; 1841, 1), but we shall here follow the investigation of sir g. g. stokes (_phil. trans._, 1848, 227), limiting ourselves, however, to the case where the retarded and unretarded beams are contiguous and of equal width. the aperture of the unretarded beam may thus be taken to be limited by x = -h, x = 0, y = -l, y= +l; and that of the beam retarded by r to be given by x = 0, x = h, y= -l, y = +l. for the former (1) § 3 gives _ _ 1 / 0 / +l / x[xi] + y[eta]\ - --------- | | sin k (at - f + -------------- )dxdy [lambda]f _/-h _/-l \ f / 2lh f k[eta]l 2f k[xi]h / [xi]h \ = - --------- · ------- sin ------- · ------ sin ------ · sin k (at - f - ----- ) (1), [lambda]f k[eta]l f k[xi]h 2f \ 2f / on integration and reduction. for the retarded stream the only difference is that we must subtract r from at, and that the limits of x are 0 and +h. we thus get for the disturbance at [xi], [eta], due to this stream 2lh f k[eta]l 2f k[xi]h / [xi]h \ - --------- · ------- sin ------- · ------ sin ------ . sin k (at - f - r + ----- ) (2). [lambda]f k[eta]l f k[xi]h 2f \ 2f / if we put for shortness [pi] for the quantity under the last circular function in (1), the expressions (1), (2) may be put under the forms u sin [tau], v sin ([tau] - [alpha]) respectively; and, if i be the intensity, i will be measured by the sum of the squares of the coefficients of sin [tau] and cos [tau] in the expression u sin[tau] + v sin([tau] - [alpha]), so that i = u2 + v2 + 2uv cos[alpha], which becomes on putting for u, v, and [alpha] their values, and putting / f k[eta]l \2 ( ------- sin ------- ) = q (3), \k[eta]l f / _ _ 4l2 [pi][xi]h | / 2[pi]r 2[pi][xi]h\ | i = q · ---------- sin2 --------- |2 + 2 cos ( -------- - ---------- ) | (4). [pi]2[xi]2 [lambda]f |_ \[lambda] [lambda]f / _| if the subject of examination be a luminous line parallel to [eta], we shall obtain what we require by integrating (4) with respect to [eta] from -[oo] to +[oo]. the constant multiplier is of no especial interest so that we may take as applicable to the image of a line _ _ 2 [pi][xi]h | / 2[pi]r 2[pi][xi]h \ | i = ----- sin2 --------- |1 + cos ( -------- - ---------- ) | (5). [xi]2 [lambda]f |_ \[lambda] [lambda]f / _| if r = 1⁄2[lambda], i vanishes at [xi]= 0; but the whole illumination, represented by _ / +[oo] | i d[xi], is independent of the value of r. if r = 0, _/-[oo] 1 2[pi][xi]h i = ----- sin2 ----------, [xi]2 [lambda]f in agreement with § 3, where a has the meaning here attached to 2h. the expression (5) gives the illumination at [xi] due to that part of the complete image whose geometrical focus is at [xi] = 0, the retardation for this component being r. since we have now to integrate for the whole illumination at a particular point o due to all the components which have their foci in its neighbourhood, we may conveniently regard o as origin. [xi] is then the co-ordinate relatively to o of any focal point o' for which the retardation is r; and the required result is obtained by simply integrating (5) with respect to [xi] from -[oo] to +[oo]. to each value of [xi] corresponds a different value of [lambda], and (in consequence of the dispersing power of the plate) of r. the variation of [lambda] may, however, be neglected in the integration, except in 2[pi]r/[lambda], where a small variation of [lambda] entails a comparatively large alteration of phase. if we write [rho] = 2[pi]r/[lambda] (6), we must regard [rho] as a function of [xi], and we may take with sufficient approximation under any ordinary circumstances [rho] = [rho]' + [=omega][xi] (7), where [rho]' denotes the value of [rho] at o, and [=omega] is a constant, which is positive when the retarding plate is held at the side on which the lue of the spectrum _is seen_. the possibility of dark bands depends upon [=omega] being positive. only in this case can cos {[rho]' + ([=omega] - 2[pi]h/[lambda]f)[xi]} retain the constant value -1 throughout the integration, and then only when [=omega] = 2[pi]h / [lambda]f (8) and cos [rho]' = -1 (9). the first of these equations is the condition for the formation of dark bands, and the second marks their situation, which is the same as that determined by the imperfect theory. the integration can be effected without much difficulty. for the first term in (5) the evaluation is effected at once by a known formula. in the second term if we observe that cos {[rho]' +([=omega] - 2[pi]h/[lambda]f)[xi]} = cos {[rho]'- g1[xi]} = cos [rho]' cos g1[xi] + sin [rho]' sin g1[xi], we see that the second part vanishes when integrated, and that the remaining integral is of the form _+[oo] / d[xi] w = | sin2 h1[xi] cos g1[xi] -----, _/-[oo] [xi]2 where h1 = [pi]h/[lambda]f, g1 = [omega] - 2[pi]h/[lambda]f (10). by differentiation with respect to g1 it may be proved that w = 0 from g1 = -[oo] to g1 = -2h1, w = 1⁄2[pi](2h1 + g1) from g1 = -2h1 to g1 = 0, w = 1⁄2[pi](2h1 - g1) from g1 = 0 to g1 = 2h1, w = 0 from g1 = 2h1 to g1 = [oo]. the integrated intensity, i', or 2[pi]h1 + 2 cos[rho]w, is thus i' = 2[pi]h1 (11), when g1 numerically exceeds 2h1; and, when g1 lies between ±2h1, i = [pi]2h1 + (2h1 - [sqrt] g12) cos[rho]' (12). it appears therefore that there are no bands at all unless [omega] lies between 0 and +4h1, and that within these limits the best bands are formed at the middle of the range when [omega] = 2h1. the formation of bands thus requires that the retarding plate be held upon the side already specified, so that [omega] be positive; and that the thickness of the plate (to which [omega] is proportional) do not exceed a certain limit, which we may call 2t0. at the best thickness t0 the bands are black, and not otherwise. the linear width of the band (e) is the increment of [xi] which alters [rho] by 2[pi], so that e = 2[pi]/[=omega] (13). with the best thickness [=omega] = 2[pi]h/[lambda]f (14), so that in this case e = [lambda]f/h (15). the bands are thus of the same width as those due to two infinitely narrow apertures coincident with the central lines of the retarded and unretarded streams, the subject of examination being itself a fine luminous line. if it be desired to see a given number of bands in the whole or in any part of the spectrum, the thickness of the retarding plate is thereby determined, independently of all other considerations. but in order that the bands may be really visible, and still more in order that they may be black, another condition must be satisfied. it is necessary that the aperture of the pupil be accommodated to the angular extent of the spectrum, or reciprocally. black bands will be too fine to be well seen unless the aperture (2h) of the pupil be somewhat contracted. one-twentieth to one-fiftieth of an inch is suitable. the aperture and the number of bands being both fixed, the condition of blackness determines the angular magnitude of a band and of the spectrum. the use of a grating is very convenient, for not only are there several spectra in view at the same time, but the dispersion can be varied continuously by sloping the grating. the slits may be cut out of tin-plate, and half covered by mica or "microscopic glass," held in position by a little cement. if a telescope be employed there is a distinction to be observed, according as the half-covered aperture is between the eye and the ocular, or in front of the object-glass. in the former case the function of the telescope is simply to increase the dispersion, and the formation of the bands is of course independent of the particular manner in which the dispersion arises. if, however, the half-covered aperture be in front of the object-glass, the phenomenon is magnified as a whole, and the desirable relation between the (unmagnified) dispersion and the aperture is the same as without the telescope. there appears to be no further advantage in the use of a telescope than the increased facility of accommodation, and for this of course a very low power suffices. the original investigation of stokes, here briefly sketched, extends also to the case where the streams are of unequal width h, k, and are separated by an interval 2g. in the case of unequal width the bands cannot be black; but if h = k, the finiteness of 2g does not preclude the formation of black bands. the theory of talbot's bands with a half-covered _circular_ aperture has been considered by h. struve (_st peters. trans._, 1883, 31, no. 1). the subject of "talbot's bands" has been treated in a very instructive manner by a. schuster (_phil. mag._, 1904), whose point of view offers the great advantage of affording an instantaneous explanation of the peculiarity noticed by brewster. a plane _pulse_, i.e. a disturbance limited to an infinitely thin slice of the medium, is supposed to fall upon a parallel grating, which again may be regarded as formed of infinitely thin wires, or infinitely narrow lines traced upon glass. the secondary pulses diverted by the ruling fall upon an object-glass as usual, and on arrival at the focus constitute a procession equally spaced in time, the interval between consecutive members depending upon the obliquity. if a retarding plate be now inserted so as to operate upon the pulses which come from one side of the grating, while leaving the remainder unaffected, we have to consider what happens at the focal point chosen. a full discussion would call for the formal application of fourier's theorem, but some conclusions of importance are almost obvious. previously to the introduction of the plate we have an effect corresponding to wave-lengths closely grouped around the principal wave-length, viz. [sigma] sin [phi], where [sigma] is the grating-interval and [phi] the obliquity, the closeness of the grouping increasing with the number of intervals. in addition to these wave-lengths there are other groups centred round the wave-lengths which are submultiples of the principal one--the overlapping spectra of the second and higher orders. suppose now that the plate is introduced so as to cover naif the aperture and that it retards those pulses which would otherwise arrive first. the consequences must depend upon the amount of the retardation. as this increases from zero, the two processions which correspond to the two halves of the aperture begin to overlap, and the overlapping gradually increases until there is almost complete superposition. the stage upon which we will fix our attention is that where the one procession bisects the intervals between the other, so that a new simple procession is constituted, containing the same number of members as before the insertion of the plate, but now spaced at intervals only half as great. it is evident that the effect at the focal point is the obliteration of the first and other spectra of odd order, so that as regards the spectrum of the first order we may consider that the two beams _interfere_. the formation of black bands is thus explained, and it requires that the plate be introduced upon one particular side, and that the amount of the retardation be adjusted to a particular value. if the retardation be too little, the overlapping of the processions is incomplete, so that besides the procession of half period there are residues of the original processions of full period. the same thing occurs if the retardation be too great. if it exceed the double of the value necessary for black bands, there is again no overlapping and consequently no interference. if the plate be introduced upon the other side, so as to retard the procession originally in arrear, there is no overlapping, whatever may be the amount of retardation. in this way the principal features of the phenomenon are accounted for, and schuster has shown further how to extend the results to spectra having their origin in prisms instead of gratings. 10. _diffraction when the source of light is not seen in focus._--the phenomena to be considered under this head are of less importance than those investigated by fraunhofer, and will be treated in less detail; but in view of their historical interest and of the ease with which many of the experiments may be tried, some account of their theory cannot be omitted. one or two examples have already attracted our attention when considering fresnel's zones, viz. the shadow of a circular disk and of a screen circularly perforated. fresnel commenced his researches with an examination of the fringes, external and internal, which accompany the shadow of a narrow opaque strip, such as a wire. as a source of light he used sunshine passing through a very small hole perforated in a metal plate, or condensed by a lens of short focus. in the absence of a heliostat the latter was the more convenient. following, unknown to himself, in the footsteps of young, he deduced the principle of interference from the circumstance that the darkness of the interior bands requires the co-operation of light from both sides of the obstacle. at first, too, he followed young in the view that the exterior bands are the result of interference between the direct light and that reflected from the edge of the obstacle, but he soon discovered that the character of the edge--e.g. whether it was the cutting edge or the back of a razor--made no material difference, and was thus led to the conclusion that the explanation of these phenomena requires nothing more than the application of huygens's principle to the unobstructed parts of the wave. in observing the bands he received them at first upon a screen of finely ground glass, upon which a magnifying lens was focused; but it soon appeared that the ground glass could be dispensed with, the diffraction pattern being viewed in the same way as the image formed by the object-glass of a telescope is viewed through the eye-piece. this simplification was attended by a great saving of light, allowing measures to be taken such as would otherwise have presented great difficulties. in theoretical investigations these problems are usually treated as of two dimensions only, everything being referred to the plane passing through the luminous point and perpendicular to the diffracting edges, supposed to be straight and parallel. in strictness this idea is appropriate only when the source is a luminous line, emitting cylindrical waves, such as might be obtained from a luminous point with the aid of a cylindrical lens. when, in order to apply huygens's principle, the wave is supposed to be broken up, the phase is the same at every element of the surface of resolution which lies upon a line perpendicular to the plane of reference, and thus the effect of the whole line, or rather infinitesimal strip, is related in a constant manner to that of the element which lies in the plane of reference, and may be considered to be represented thereby. the same method of representation is applicable to spherical waves, issuing from a _point_, if the radius of curvature be large; for, although there is variation of phase along the length of the infinitesimal strip, the whole effect depends practically upon that of the central parts where the phase is sensibly constant.[10] [illustration: fig. 17.] in fig. 17 apq is the arc of the circle representative of the wave-front of resolution, the centre being at o, and the radius qa being equal to a. b is the point at which the effect is required, distant a + b from o, so that ab = b, ap = s, pq = ds. taking as the standard phase that of the secondary wave from a, we may represent the effect of pq by /t [delta] \ cos 2[pi] ( - - -------- )·ds, \r [lambda]/ where [delta] = bp - ap is the retardation at b of the wave from p relatively to that from a. now [delta] = (a + b) s2/2ab (1), so that, if we write 2[pi][delta] = [pi](a + b)s2 [pi]v2 ------------ --------------- = ------ (2), [lambda] ab[lambda] 2 the effect at b is _ _ /ab[lambda]\1⁄2 / 2[pi]t / 2[pi]t / \ ( ---------- ) ( cos ------ | cos 1⁄2[pi]v2·dv + sin ------ | sin 1⁄2[pi]v2·dv ) (3), \2(a + b) / \ [tau] _/ [tau] _/ / the limits of integration depending upon the disposition of the diffracting edges. when a, b, [lambda] are regarded as constant, the first factor may be omitted,--as indeed should be done for consistency's sake, inasmuch as other factors of the same nature have been omitted already. the intensity i2, the quantity with which we are principally concerned, may thus be expressed _ _ / / \2 / / \2 i2= ( | cos 1⁄2[pi]v2·dv ) + ( | sin 1⁄2[pi]v2·dv ) (4). \ _/ / \ _/ / these integrals, taken from v = 0, are known as fresnel's integrals; we will denote them by c and s, so that _ _ / v / v c = | cos 1⁄2[pi]v2·dv, s = | cos 1⁄2[pi]v2·dv (5). _/0 _/0 when the upper limit is infinity, so that the limits correspond to the inclusion of half the primary wave, c and s are both equal to 1⁄2, by a known formula; and on account of the rapid fluctuation of sign the parts of the range beyond very moderate values of v contribute but little to the result. ascending series for c and s were given by k. w. knockenhauer, and are readily investigated. integrating by parts, we find _v _v / i·1⁄2[pi]v2 i·1⁄2[pi]v2 1 / i·1⁄2[pi]v2 c + is = | e dv = e · v - - i[pi] | e dv3; _/0 3 _/0 and, by continuing this process, i.1⁄2[pi]v2 / i[pi] i[pi] i[pi] i[pi] i[pi] i[pi] \ c + is = e ( v - ----- v3 + ----- ----- v^5 - ----- ----- ----- v^7 + ... ). \ 3 3 5 3 5 7 / by separation of real and imaginary parts, c = m cos 1⁄2[pi]v2 - n sin 1⁄2[pi]v2 \ s = m sin 1⁄2[pi]v2 - n cos 1⁄2[pi]v2 / (6) where v [pi]2v^5 [pi]^4v^9 m = - - --------- + --------- - ... (7) 1 3·5 3·5·7·9 [pi]v3 [pi]^3v^7 [pi]^5v^11 n = ------ - --------- + ------------ ... (8) 1·3 1·3·5·7 1·3·5·7·9·11 these series are convergent for all values of v, but are practically useful only when v is small. expressions suitable for discussion when v is large were obtained by l. p. gilbert (_mem. cour. de l'acad. de bruxelles_, 31, p. 1). taking 1⁄2[pi]v2 = u (9), we may write _ 1 /u e^iu du c + is = ------------- | -------- (10). [sqrt](2[pi]) _/0 [sqrt] u again, by a known formula, _[oo] 1 1 / e^-ux dx -------- = ---------- | -------- (11). [sqrt] u [sqrt][pi] _/0 [sqrt]x substituting this in (10), and inverting the order of integration, we get _[oo] _u 1 / dx / e^u(i - x) c + is = ------- | -------- | ----------- dx [sqrt]2 _/0 [sqrt] x _/0 [sqrt]x _[oo] 1 / dx e^u(i - x) - 1 = ------- | -------- -------------- dx (12). [sqrt]2 _/0 [sqrt] x i - x thus, if we take _[oo] 1 / e^-ux [sqrt](x)·dx g = ----------- | ------------------, [pi][sqrt]2 _/0 1 + x2 _[oo] 1 / e^-ux dx h = ----------- | ------------------ (13). [pi][sqrt]2 _/ [sqrt]x · (1 + x2) 0 c = 1⁄2 - g cos u + h sin u, s = 1⁄2 - g sin u - h cos u (14). the constant parts in (14), viz. 1⁄2, may be determined by direct integration of (12), or from the observation that by their constitution g and h vanish when u = [oo], coupled with the fact that c and s then assume the value 1⁄2. comparing the expressions for c, s in terms of m, n, and in terms of g, h, we find that g = 1⁄2 (cos u + sin u) - m, h = 1⁄2 (cos u - sin u) + n (15), formulae which may be utilized for the calculation of g, h when u (or v) is small. for example, when u = 0, m = 0, n = 0, and consequently g = h = 1⁄2. descending series of the semi-convergent class, available for numerical calculation when u is moderately large, can be obtained from (12) by writing x = uy, and expanding the denominator in powers of y. the integration of the several terms may then be effected by the formula _ [oo] / -y q-1⁄2 | e y dy = [gamma](q + 1⁄2) = (q - 1⁄2)(q - 3/2) ... 1⁄2[sqrt][pi]; _/0 and we get in terms of v 1 1·3·5 1·3·5·9 g = ------- - ---------- + ----------- - (16), [pi]2v3 [pi]^4 v^7 [pi]^6 v^11 1 1·3 1·3·5·7 h = ----- - --------- + ---------- - (17). [pi]v [pi]3 v^5 [pi]^5 v^9 the corresponding values of c and s were originally derived by a. l. cauchy, without the use of gilbert's integrals, by direct integration by parts. from the series for g and h just obtained it is easy to verify that dh dg -- = - [pi]vg, -- = [pi]vh - 1 (18). dv dv we now proceed to consider more particularly the distribution of light upon a screen pbq near the shadow of a straight edge a. at a point p within the geometrical shadow of the obstacle, the half of the wave to the right of c (fig. 18), the nearest point on the wave-front, is wholly intercepted, and on the left the integration is to be taken from s = ca to s = [oo]. if v be the value of v corresponding to ca, viz. / / 2(a + b) \ v= / ( ---------- )·ca, (19), \/ \ ab[lambda] / we may write _[oo] _[oo] / / \2 / / \2 i2 = ( | cos 1⁄2[pi]v2·dv ) + ( | sin 1⁄2[pi]v2·dv ) (20), \ _/v / \ _/v / or, according to our previous notation, i2 = (1⁄2 - cv)2 + (1⁄2 - sv)2 = g2 + h2 (21). now in the integrals represented by g and h every element diminishes as v increases from zero. hence, as ca increases, viz. as the point p is more and more deeply immersed in the shadow, the illumination _continuously_ decreases, and that without limit. it has long been known from observation that there are no bands on the interior side of the shadow of the edge. [illustration: fig. 18.] the law of diminution when v is moderately large is easily expressed with the aid of the series (16), (17) for g, h. we have ultimately g = 0, h = ([pi]v)^-1, so that i2 = 1/[pi]2v2, or the illumination is inversely as the square of the distance from the shadow of the edge. for a point q outside the shadow the integration extends over _more_ than half the primary wave. the intensity may be expressed by i2 = (1⁄2 + cv)2 + (1⁄2 + sv)2 (22); and the maxima and minima occur when dc ds (1⁄2 + c_v) -- + (1⁄2 + s_v) -- = 0, dv dv whence sin 1⁄2[pi]v2 + cos 1⁄2[pi]v2 = g (23). when v = 0, viz. at the edge of the shadow, i2 = 1⁄2; when v = [oo], i2 = 2, on the scale adopted. the latter is the intensity due to the uninterrupted wave. the quadrupling of the intensity in passing outwards from the edge of the shadow is, however, accompanied by fluctuations giving rise to bright and dark bands. the position of these bands determined by (23) may be very simply expressed when v is large, for then sensibly g = 0, and 1⁄2[pi]v2 = 3⁄4[pi] + n[pi] (24), n being an integer. in terms of [delta], we have from (2) [delta] = (3/8 + 1⁄2n)[lambda] (25). the first maximum in fact occurs when [delta] = 3/8[lambda] -.0046[lambda], and the first minimum when [delta] = 7/8[lambda] -.0016[lambda], the corrections being readily obtainable from a table of g by substitution of the approximate value of v. the position of q corresponding to a given value of v, that is, to a band of given order, is by (19) a + b / / b[lambda](a + b) \ bq = ----- ad = v / ( ----------------- ) (26). a \/ \ 2a / by means of this expression we may trace the locus of a band of given order as b varies. with sufficient approximation we may regard bq and b as rectangular co-ordinates of q. denoting them by x, y, so that ab is axis of y and a perpendicular through a the axis of x, and rationalizing (26), we have 2ax2 - v2[lambda]y2 - v2a[lambda]y = 0, which represents a hyperbola with vertices at o and a. from (24), (26) we see that the width of the bands is of the order [sqrt] {b[lambda](a + b)/a}. from this we may infer the limitation upon the width of the source of light, in order that the bands may be properly formed. if [omega] be the apparent magnitude of the source seen from a, [omega]b should be much smaller than the above quantity, or [omega] < [sqrt] {[lambda](a + b)/ab} (27). if a be very great in relation to b, the condition becomes [omega] < [sqrt] ([lambda]/b) (28). so that if b is to be moderately great (1 metre), the apparent magnitude of the sun must be greatly reduced before it can be used as a source. the values of v for the maxima and minima of intensity, and the magnitudes of the latter, were calculated by fresnel. an extract from his results is given in the accompanying table. +--------------------+----------+------------+ | | v | i2 | +--------------------+----------+------------+ | first maximum | 1.2172 | 2.7413 | | first minimum | 1.8726 | 1.5570 | | second maximum | 2.3449 | 2.3990 | | second minimum | 2.7392 | 1.6867 | | third maximum. | 3.0820 | 2.3022 | | third minimum | 3.3913 | 1.7440 | +--------------------+----------+------------+ a very thorough investigation of this and other related questions, accompanied by fully worked-out tables of the functions concerned, will be found in a paper by e. lommel (_abh. bayer. akad. d. wiss._ ii. ci., 15, bd., iii. abth., 1886). when the functions c and s have once been calculated, the discussion of various diffraction problems is much facilitated by the idea, due to m. a. cornu (_journ. de phys._, 1874, 3, p. 1; a similar suggestion was made independently by g. f. fitzgerald), of exhibiting as a curve the relationship between c and s, considered as the rectangular co-ordinates (x, y) of a point. such a curve is shown in fig. 19, where, according to the definition (5) of c, s, _ v _ v / / x = | cos 1⁄2[pi]v2·dv, y = | sin 1⁄2[pi]v2·dv (29). _/0 _/0 the origin of co-ordinates o corresponds to v = 0; and the asymptotic points j, j', round which the curve revolves in an ever-closing spiral, correspond to v = ±[oo]. the intrinsic equation, expressing the relation between the arc [sigma] (measured from o) and the inclination [phi] of the tangent at any points to the axis of x, assumes a very simple form. for dx = cos 1⁄2[pi]v2·dv, dy = sin 1⁄2[pi]v2·dv; so that _ / [sigma] = | [sqrt] (dx2 + dy2) = v, (30), _/ [phi] = tan^-1 (dy/dx) = 1⁄2[pi]v2 (31). accordingly, [phi] = 1⁄2[pi][sigma]2 (32); and for the curvature, d[phi]/d[sigma] = [pi][sigma] (33). cornu remarks that this equation suffices to determine the general character of the curve. for the osculating circle at any point includes the whole of the curve which lies beyond; and the successive convolutions envelop one another without intersection. [illustration: fig. 19.] the utility of the curve depends upon the fact that the elements of arc represent, in amplitude and phase, the component vibrations due to the corresponding portions of the primary wave-front. for by (30) d[sigma] = dv, and by (2) dv is proportional to ds. moreover by (2) and (31) the retardation of phase of the elementary vibration from pq (fig. 17) is 2[pi][delta]/[lambda], or [phi]. hence, in accordance with the rule for compounding vector quantities, the resultant vibration at b, due to any finite part of the primary wave, is represented in amplitude and phase by the chord joining the extremities of the corresponding arc ([sigma]2 - [sigma]1). in applying the curve in special cases of diffraction to exhibit the effect at any point p (fig. 18) the centre of the curve o is to be considered to correspond to that point c of the primary wave-front which lies nearest to p. the operative part, or parts, of the curve are of course those which represent the unobstructed portions of the primary wave. let us reconsider, following cornu, the diffraction of a screen unlimited on one side, and on the other terminated by a straight edge. on the illuminated side, at a distance from the shadow, the vibration is represented by jj'. the co-ordinates oi j, j' being (1⁄2, 1⁄2), (-1⁄2, -1⁄2), i2 is 2; and the phase is 1/8 period in arrear of that of the element at o. as the point under contemplation is supposed to approach the shadow, the vibration is represented by the chord drawn from j to a point on the other half of the curve, which travels inwards from j' towards o. the amplitude is thus subject to fluctuations, which increase as the shadow is approached. at the point o the intensity is one-quarter of that of the entire wave, and after this point is passed, that is, when we have entered the geometrical shadow, the intensity falls off gradually to zero, _without fluctuations_. the whole progress of the phenomenon is thus exhibited to the eye in a very instructive manner. we will next suppose that the light is transmitted by a slit, and inquire what is the effect of varying the width of the slit upon the illumination at the projection of its centre. under these circumstances the arc to be considered is bisected at o, and its length is proportional to the width of the slit. it is easy to see that the length of the chord (which passes in all cases through o) increases to a maximum near the place where the phase-retardation is 3/8 of a period, then diminishes to a minimum when the retardation is about 7/8 of a period, and so on. if the slit is of constant width and we require the illumination at various points on the screen behind it, we must regard the arc of the curve as of _constant length_. the intensity is then, as always, represented by the square of the length of the chord. if the slit be narrow, so that the arc is short, the intensity is constant over a wide range, and does not fall off to an important extent until the discrepancy of the extreme phases reaches about a quarter of a period. we have hitherto supposed that the shadow of a diffracting obstacle is received upon a diffusing screen, or, which comes to nearly the same thing, is observed with an eye-piece. if the eye, provided if necessary with a perforated plate in order to reduce the aperture, be situated inside the shadow at a place where the illumination is still sensible, and be focused upon the diffracting edge, the light which it receives will appear to come from the neighbourhood of the edge, and will present the effect of a silver lining. this is doubtless the explanation of a "pretty optical phenomenon, seen in switzerland, when the sun rises from behind distant trees standing on the summit of a mountain."[11] ii. _dynamical theory of diffraction._--the explanation of diffraction phenomena given by fresnel and his followers is independent of special views as to the nature of the aether, at least in its main features; for in the absence of a more complete foundation it is impossible to treat rigorously the mode of action of a solid obstacle such as a screen. but, without entering upon matters of this kind, we may inquire in what manner a primary wave may be resolved into elementary secondary waves, and in particular as to the law of intensity and polarization in a secondary wave as dependent upon its direction of propagation, and upon the character as regards polarization of the primary wave. this question was treated by stokes in his "dynamical theory of diffraction" (_camb. phil. trans._, 1849) on the basis of the elastic solid theory. let x, y, z be the co-ordinates of any particle of the medium in its natural state, and [chi], [eta], [zeta] the displacements of the same particle at the end of time t, measured in the directions of the three axes respectively. then the first of the equations of motion may be put under the form d2[xi] /d2[xi] d2[xi] d2[xi]\ d2 /d2[xi] d2[eta] d2[zeta]\ ------ = b2( ------ + ------ + ------ ) + (a2 - b2)--( ------ + ------- + -------- ), dt2 \ dx2 dy2 dz2 / dx \ dx2 dy2 dz2 / where a2 and b2 denote the two arbitrary constants. put for shortness d2[xi] d2[eta] d2[zeta] ------ + ------- + -------- = [delta] (1), dx2 dy2 dz2 and represent by [delta]2[chi] the quantity multiplied by b2. according to this notation, the three equations of motion are d2[xi] d[delta] \ ------ = b2[delta]2[xi] + (a2 - b2) -------- | dt2 dx | | d2[eta] d[delta] | ------- = b2[delta]2[eta] + (a2 - b2) -------- > (2). dt2 dy | | d2[zeta] d[delta] | -------- = b2[delta]2[zeta] + (a2 - b2) -------- | dt2 dz / it is to be observed that s denotes the dilatation of volume of the element situated at (x, y, z). in the limiting case in which the medium is regarded as absolutely incompressible [delta] vanishes; but, in order that equations (2) may preserve their generality, we must suppose a at the same time to become infinite, and replace a2[delta] by a new function of the co-ordinates. these equations simplify very much in their application to plane waves. if the ray be parallel to ox, and the direction of vibration parallel to oz, we have [xi] = 0, [eta] = 0, while [zeta] is a function of x and t only. equation (1) and the first pair of equations (2) are thus satisfied identically. the third equation gives d2[zeta] d2[zeta] -------- = -------- (3), dt2 dx2 of which the solution is [zeta] = f(bt - x) (4), where f is an arbitrary function. the question as to the law of the secondary waves is thus answered by stokes. "let [xi] = 0, [eta] = 0, [zeta] = f(bt-x) be the displacements corresponding to the incident light; let o1 be any point in the plane p (of the wave-front), ds an element of that plane adjacent to o1, and consider the disturbance due to that portion only of the incident disturbance which passes continually across ds. let o be any point in the medium situated at a distance from the point o1 which is large in comparison with the length of a wave; let o1o = r, and let this line make an angle [theta] with the direction of propagation of the incident light, or the axis of x, and [phi] with the direction of vibration, or axis of z. then the displacement at o will take place in a direction perpendicular to o1o, and lying in the plane zo1o; and, if [zeta]' be the displacement at o, reckoned positive in the direction nearest to that in which the incident vibrations are reckoned positive, ds [zeta]' = ------ ( 1 + cos[theta]) sin[phi] f'(bt - r). 4[pi]r in particular, if 2[pi] f(bt - x) = c sin -------- (bt - x) (5), [lambda] we shall have cds 2[pi] [zeta]' = ---------- (1 + cos[theta]) sin[phi]cos -------- (bt - r) (6)." 2[lambda]r [lambda] it is then verified that, after integration with respect to ds, (6) gives the same disturbance as if the primary wave had been supposed to pass on unbroken. the occurrence of sin [phi] as a factor in (6) shows that the relative intensities of the primary light and of that diffracted in the direction [theta] depend upon the condition of the former as regards polarization. if the direction of primary vibration be perpendicular to the plane of diffraction (containing both primary and secondary rays), sin [phi] = 1; but, if the primary vibration be in the plane of diffraction, sin [phi] = cos [theta]. this result was employed by stokes as a criterion of the direction of vibration; and his experiments, conducted with gratings, led him to the conclusion that the vibrations of polarized light are executed in a direction _perpendicular_ to the plane of polarization. the factor (1 + cos [theta]) shows in what manner the secondary disturbance depends upon the direction in which it is propagated with respect to the front of the primary wave. if, as suffices for all practical purposes, we limit the application of the formulae to points in advance of the plane at which the wave is supposed to be broken up, we may use simpler methods of resolution than that above considered. it appears indeed that the purely mathematical question has no definite answer. in illustration of this the analogous problem for sound may be referred to. imagine a flexible lamina to be introduced so as to coincide with the plane at which resolution is to be effected. the introduction of the lamina (supposed to be devoid of inertia) will make no difference to the propagation of plane parallel sonorous waves through the position which it occupies. at every point the motion of the lamina will be the same as would have occurred in its absence, the pressure of the waves impinging from behind being just what is required to generate the waves in front. now it is evident that the aerial motion in front of the lamina is determined by what happens at the lamina without regard to the cause of the motion there existing. whether the necessary forces are due to aerial pressures acting on the rear, or to forces directly impressed from without, is a matter of indifference. the conception of the lamina leads immediately to two schemes, according to which a primary wave may be supposed to be broken up. in the first of these the element ds, the effect of which is to be estimated, is supposed to execute its actual motion, while every other element of the plane lamina is maintained at rest. the resulting aerial motion in front is readily calculated (see rayleigh, _theory of sound_, § 278); it is symmetrical with respect to the origin, i.e. independent of [theta]. when the secondary disturbance thus obtained is integrated with respect to ds over the entire plane of the lamina, the result is necessarily the same as would have been obtained had the primary wave been supposed to pass on without resolution, for this is precisely the motion generated when every element of the lamina vibrates with a common motion, equal to that attributed to ds. the only assumption here involved is the evidently legitimate one that, when two systems of variously distributed motion at the lamina are superposed, the corresponding motions in front are superposed also. the method of resolution just described is the simplest, but it is only one of an indefinite number that might be proposed, and which are all equally legitimate, so long as the question is regarded as a merely mathematical one, without reference to the physical properties of actual screens. if, instead of supposing the _motion_ at ds to be that of the primary wave, and to be zero elsewhere, we suppose the _force_ operative over the element ds of the lamina to be that corresponding to the primary wave, and to vanish elsewhere, we obtain a secondary wave following quite a different law. in this case the motion in different directions varies as cos[theta], vanishing at right angles to the direction of propagation of the primary wave. here again, on integration over the entire lamina, the aggregate effect of the secondary waves is necessarily the same as that of the primary. in order to apply these ideas to the investigation of the secondary wave of light, we require the solution of a problem, first treated by stokes, viz. the determination of the motion in an infinitely extended elastic solid due to a locally applied periodic force. if we suppose that the force impressed upon the element of mass d dx dy dz is dz dx dy dz, being everywhere parallel to the axis of z, the only change required in our equations (1), (2) is the addition of the term z to the second member of the third equation (2). in the forced vibration, now under consideration, z, and the quantities [xi], [eta], [zeta], [delta] expressing the resulting motion, are to be supposed proportional to e^int, where i = [sqrt](-1), and n = 2[pi]/[tau], [tau] being the periodic time. under these circumstances the double differentiation with respect to t of any quantity is equivalent to multiplication by the factor -n2, and thus our equations take the form d[delta] \ (b2[delta]2 + n2)[xi] + (a2 - b2) -------- = 0 | dx | | d[delta] | (b2[delta]2 + n2)[eta] + (a2 - b2) -------- = 0 > (7). dx | | d[delta] | (b2[delta]2 + n2)[zeta] + (a2 - b2) -------- = -z | dx / it will now be convenient to introduce the quantities.[=omega]1, [=omega]2, [=omega]3 which express the _rotations_ of the elements of the medium round axes parallel to those of co-ordinates, in accordance with the equations d[xi] d[eta] d[eta] d[zeta] [=omega]3 = ----- - ------, [=omega]1 = ------ - -------, dy dx' dz dy d[zeta] d[xi] [=omega]2 = ------- - ----- (8). dx dz in terms of these we obtain from (7), by differentiation and subtraction, (b2[delta]2 + n2) [=omega]3 = 0 \ (b2[delta]2 + n2) [=omega]1 = dz/dy > (9). (b2[delta]2 + n2) [=omega]2 = -dz/dx / the first of equations (9) gives [=omega]3 = 0 (10). for =[omega]1, we have _ _ _ -ikr 1 / / / dz e [=omega]1 = ------- | | | -- ----- dx dy dz (11), 4[pi]b2 _/_/_/ dy r where r is the distance between the element dx dy dz and the point where [=omega]1 is estimated, and k = n/b = 2[pi]/[lambda] (12), [lambda] being the wave-length. (this solution may be verified in the same manner as poisson's theorem, in which k = 0.) we will now introduce the supposition that the force z acts only within a small space of volume t, situated at (x, y, z), and for simplicity suppose that it is at the origin of co-ordinates that the rotations are to be estimated. integrating by parts in (11), we get _ -ikr _ _ _ / e dz | ze^-ikr | / d / e^-ikr\ | ------ -- dy = | ------- | - | z -- ( ------- ) dy, _/ r dy |_ r _| _/ dy \ r / in which the integrated terms at the limits vanish, z being finite only within the region t. thus _ _ _ -ikr 1 / / / d /e^ \ [=omega]1 = ------- | | | z -- ( -------- ) dx dy dz. 4[pi]b2 _/_/_/ dy \ r / since the dimensions of t are supposed to be very small in comparison with [lambda], the factor d/dy (e^-ikr / r) is sensibly constant; so that, if z stand for the mean value of z over the volume t, we may write tz y d / e^-ikr \ [=omega]1 = ------- · - · -- ( ------ ) (13). 4[pi]b2 r dr \ r / in like manner we find tz x d / e^-ikr \ [=omega]2 = ------ · - · -- ( ------- ) (14). 4[pi]b2 r dr \ r / from (10), (13), (14) we see that, as might have been expected, the rotation at any point is about an axis perpendicular both to the direction of the force and to the line joining the point to the source of disturbance. if the resultant rotation be [omega], we have tz [sqrt](x2 + y2) d /e^-ikr\ [=omega] = ------- · --------------- · -- ( ------ ) = 4[pi]b2 r dr \ r / tz sin[phi] d /e^-ikr\ = ----------- -- ( ------ ), 4[pi]b2 dr \ r / [phi] denoting the angle between r and z. in differentiating e^(-ikr)/r with respect to r, we may neglect the term divided by r2 as altogether insensible, kr being an exceedingly great quantity at any moderate distance from the origin of disturbance. thus -ik·tz sin[phi] /e^-ikr\ [=omega] = --------------- · ( ------ ) (15), 4[pi]b2 \ r / which completely determines the rotation at any point. for a disturbing force of given integral magnitude it is seen to be everywhere about an axis perpendicular to r and the direction of the force, and in magnitude dependent only upon the angle ([phi]) between these two directions and upon the distance (r). the intensity of light is, however, more usually expressed in terms of the actual displacement in the plane of the wave. this displacement, which we may denote by [zeta]', is in the plane containing z and r, and perpendicular to the latter. its connexion with [=omega]is expressed by [=omega] = d[zeta]'/dr; so that tz sin [phi] /e^-ikr\ [zeta]' = ----------- · ( ------ ) (16), 4[pi]b2 \ r / where the factor e^int is restored. retaining only the real part of (16), we find, as the result of a local application of force equal to dtz cos nt (17), the disturbance expressed by tz sin [phi] /cos(nt - kr)\ [zeta]' = ------------ · ( ------------ ) (18). 4[pi]b2 \ r / the occurrence of sin [phi] shows that there is no disturbance radiated in the direction of the force, a feature which might have been anticipated from considerations of symmetry. we will now apply (18) to the investigation of a law of secondary disturbance, when a primary wave [zeta] = sin(nt - kx) (19) is supposed to be broken up in passing the plane x = 0. the first step is to calculate the force which represents the reaction between the parts of the medium separated by x = 0. the force operative upon the positive half is parallel to oz, and of amount per unit of area equal to -b2d d[zeta]/dx = b2kd cos nt; and to this force acting over the whole of the plane the actual motion on the positive side may be conceived to be due. the secondary disturbance corresponding to the element ds of the plane may be supposed to be that caused by a force of the above magnitude acting over ds and vanishing elsewhere; and it only remains to examine what the result of such a force would be. now it is evident that the force in question, supposed to act upon the positive half only of the medium, produces just double of the effect that would be caused by the same force if the medium were undivided, and on the latter supposition (being also localized at a point) it comes under the head already considered. according to (18), the effect of the force acting at ds parallel to oz, and of amount equal to 2b2kd ds cos nt, will be a disturbance ds sin [phi] [zeta]' = ------------ cos(nt - kr) (20), [lambda]r regard being had to (12). this therefore expresses the secondary disturbance at a distance r and in a direction making an angle [phi] with oz (the direction of primary vibration) due to the element ds of the wave-front. the proportionality of the secondary disturbance to sin [phi] is common to the present law and to that given by stokes, but here there is no dependence upon the angle [theta] between the primary and secondary rays. the occurrence of the factor [lambda]r^-1, and the necessity of supposing the phase of the secondary wave accelerated by a quarter of an undulation, were first established by archibald smith, as the result of a comparison between the primary wave, supposed to pass on without resolution, and the integrated effect of all the secondary waves (§ 2). the occurrence of factors such as sin [phi], or 1⁄2(1 + cos [theta]), in the expression of the secondary wave has no influence upon the result of the integration, the effects of all the elements for which the factors differ appreciably from unity being destroyed by mutual interference. the choice between various methods of resolution, all mathematically admissible, would be guided by physical considerations respecting the mode of action of obstacles. thus, to refer again to the acoustical analogue in which plane waves are incident upon a perforated rigid screen, the circumstances of the case are best represented by the first method of resolution, leading to symmetrical secondary waves, in which the normal motion is supposed to be zero over the unperforated parts. indeed, if the aperture is very small, this method gives the correct result, save as to a constant factor. in like manner our present law (20) would apply to the kind of obstruction that would be caused by an actual physical division of the elastic medium, extending over the whole of the area supposed to be occupied by the intercepting screen, but of course not extending to the parts supposed to be perforated. on the electromagnetic theory, the problem of diffraction becomes definite when the properties of the obstacle are laid down. the simplest supposition is that the material composing the obstacle is perfectly conducting, i.e. perfectly reflecting. on this basis a. j. w. sommerfeld (_math. ann._, 1895, 47, p. 317), with great mathematical skill, has solved the problem of the shadow thrown by a semi-infinite plane screen. a simplified exposition has been given by horace lamb (_proc. lond. math. soc._, 1906, 4, p. 190). it appears that fresnel's results, although based on an imperfect theory, require only insignificant corrections. problems not limited to two dimensions, such for example as the shadow of a circular disk, present great difficulties, and have not hitherto been treated by a rigorous method; but there is no reason to suppose that fresnel's results would be departed from materially. (r.) footnotes: [1] the descending series for j0(z) appears to have been first given by sir w. hamilton in a memoir on "fluctuating functions," _roy. irish trans._, 1840. [2] airy, loc. cit. "thus the magnitude of the central spot is diminished, and the brightness of the rings increased, by covering the central parts of the object-glass." [3] _"man kann daraus schliessen, was moglicher weise durch mikroskope noch zu sehen ist. ein mikroskopischer gegenstand z. b, dessen durchmesser = ([lambda]) ist, und der aus zwei theilen besteht, kann nicht mehr als aus zwei theilen bestehend erkannt werden. dieses zeigt uns eine grenze des sehvermogens durch mikroskope"_ (_gilbert's ann._ 74, 337). lord rayleigh has recorded that he was himself convinced by fraunhofer's reasoning at a date antecedent to the writings of helmholtz and abbe. [4] the last sentence is repeated from the writer's article "wave theory" in the 9th edition of this work, but a. a. michelson's ingenious echelon grating constitutes a realization in an unexpected manner of what was thought to be impracticable.--[r.] [5] compare also f. f. lippich, _pogg. ann._ cxxxix. p. 465, 1870; rayleigh, _nature_ (october 2, 1873). [6] the power of a grating to construct light of nearly definite wave-length is well illustrated by young's comparison with the production of a musical note by reflection of a sudden sound from a row of palings. the objection raised by herschel (_light_, § 703) to this comparison depends on a misconception. [7] it must not be supposed that errors of this order of magnitude are unobjectionable in all cases. the position of the middle of the bright band representative of a mathematical line can be fixed with a spider-line micrometer within a small fraction of the width of the band, just as the accuracy of astronomical observations far transcends the separating power of the instrument. [8] "in the same way we may conclude that in flat gratings any departure from a straight line has the effect of causing the dust in the slit and the spectrum to have different foci--a fact sometimes observed." (rowland, "on concave gratings for optical purposes," _phil. mag._, september 1883). [9] on account of inequalities in the atmosphere giving a variable refraction, the light from a star would be irregularly distributed over a screen. the experiment is easily made on a laboratory scale, with a small source of light, the rays from which, in their course towards a rather distant screen, are disturbed by the neighbourhood of a heated body. at a moment when the eye, or object-glass of a telescope, occupies a dark position, the star vanishes. a fraction of a second later the aperture occupies a bright place, and the star reappears. according to this view the chromatic effects depend entirely upon atmospheric dispersion. [10] in experiment a line of light is sometimes substituted for a point in order to increase the illumination. the various parts of the line are here _independent_ sources, and should be treated accordingly. to assume a cylindrical form of primary wave would be justifiable only when there is synchronism among the secondary waves issuing from the various centres. [11] h. necker (_phil. mag._, november 1832); fox talbot (_phil. mag._, june 1833). "when the sun is about to emerge ... every branch and leaf is lighted up with a silvery lustre of indescribable beauty.... the birds, as mr necker very truly describes, appear like flying brilliant sparks." talbot ascribes the appearance to diffraction; and he recommends the use of a telescope. diffusion (from the lat. _diffundere; dis-_, asunder, and _fundere_, to pour out), in general, a spreading out, scattering or circulation; in physics the term is applied to a special phenomenon, treated below. 1. _general description._--when two different substances are placed in contact with each other they sometimes remain separate, but in many cases a gradual mixing takes place. in the case where both the substances are gases the process of mixing continues until the result is a uniform mixture. in other cases the proportions in which two different substances can mix lie between certain fixed limits, but the mixture is distinguished from a chemical compound by the fact that between these limits the composition of the mixture is capable of continuous variation, while in chemical compounds, the proportions of the different constituents can only have a discrete series of numerical values, each different ratio representing a different compound. if we take, for example, air and water in the presence of each other, air will become dissolved in the water, and water will evaporate into the air, and the proportions of either constituent absorbed by the other will vary continuously. but a limit will come when the air will absorb no more water, and the water will absorb no more air, and throughout the change a definite surface of separation will exist between the liquid and the gaseous parts. when no surface of separation ever exists between two substances they must necessarily be capable of mixing in all proportions. if they are not capable of mixing in all proportions a discontinuous change must occur somewhere between the regions where the substances are still unmixed, thus giving rise to a surface of separation. the phenomena of mixing thus involves the following processes:--(1) a motion of the substances relative to one another throughout a definite _region_ of space in which mixing is taking place. this relative motion is called "diffusion." (2) the passage of portions of the mixing substances across the _surface_ of separation when such a surface exists. these surface actions are described under various terms such as solution, evaporation, condensation and so forth. for example, when a soluble salt is placed in a liquid, the process which occurs at the surface of the salt is called "solution," but the salt which enters the liquid by solution is transported from the surface into the interior of the liquid by "diffusion." diffusion may take place in solids, that is, in regions occupied by matter which continues to exhibit the properties of the solid state. thus if two liquids which can mix are separated by a membrane or partition, the mixing may take place through the membrane. if a solution of salt is separated from pure water by a sheet of parchment, part of the salt will pass through the parchment into the water. if water and glycerin are separated in this way most of the water will pass into the glycerin and a little glycerin will pass through in the opposite direction, a property frequently used by microscopists for the purpose of gradually transferring minute algae from water into glycerin. a still more interesting series of examples is afforded by the passage of gases through partitions of metal, notably the passage of hydrogen through platinum and palladium at high temperatures. when the process is considered with reference to a membrane or partition taken as a whole, the passage of a substance from one side to the other is commonly known as "osmosis" or "transpiration" (see solution), but what occurs in the material of the membrane itself is correctly described as diffusion. simple cases of diffusion are easily observed qualitatively. if a solution of a coloured salt is carefully introduced by a funnel into the bottom of a jar containing water, the two portions will at first be fairly well defined, but if the mixture can exist in all proportions, the surface of separation will gradually disappear; and the rise of the colour into the upper part and its gradual weakening in the lower part, may be watched for days, weeks or even longer intervals. the diffusion of a strong aniline colouring matter into the interior of gelatine is easily observed, and is commonly seen in copying apparatus. diffusion of gases may be shown to exist by taking glass jars containing vapours of hydrochloric acid and ammonia, and placing them in communication with the heavier gas downmost. the precipitation of ammonium chloride shows that diffusion exists, though the chemical action prevents this example from forming a typical case of diffusion. again, when a film of canada balsam is enclosed between glass plates, the disappearance during a few weeks of small air bubbles enclosed in the balsam can be watched under the microscope. in fluid media, whether liquids or gases, the process of mixing is greatly accelerated by stirring or agitating the fluids, and liquids which might take years to mix if left to themselves can thus be mixed in a few seconds. it is necessary to carefully distinguish the effects of agitation from those of diffusion proper. by shaking up two liquids which do not mix we split them up into a large number of different portions, and so greatly increase the area of the surface of separation, besides decreasing the thicknesses of the various portions. but even when we produce the appearance of a uniform turbid mixture, the small portions remain quite distinct. if however the fluids can really mix, the final process must in every case depend on diffusion, and all we do by shaking is to increase the sectional area, and decrease the thickness of the diffusing portions, thus rendering the completion of the operation more rapid. if a gas is shaken up in a liquid the process of absorption of the bubbles is also accelerated by capillary action, as occurs in an ordinary sparklet bottle. to state the matter precisely, however finely two fluids have been subdivided by agitation, the molecular constitution of the different portions remains unchanged. the ultimate process by which the individual molecules of two different substances become mixed, producing finally a homogeneous mixture, is in every case diffusion. in other words, diffusion is that relative motion of the molecules of two different substances by which the proportions of the molecules in any region containing a finite number of molecules are changed. in order, therefore, to make accurate observations of diffusion in fluids it is necessary to guard against any cause which may set up currents; and in some cases this is exceedingly difficult. thus, if gas is absorbed at the upper surface of a liquid, and if the gaseous solution is heavier than the pure liquid, currents may be set up, and a steady state of diffusion may cease to exist. this has been tested experimentally by c. g. von hufner and w. e. adney. the same thing may happen when a gas is evolved into a liquid at the surface of a solid even if no bubbles are formed; thus if pieces of aluminium are placed in caustic soda, the currents set up by the evolution of hydrogen are sufficient to set the aluminium pieces in motion, and it is probable that the motions of the diatomaceae are similarly caused by the evolution of oxygen. in some pairs of substances diffusion may take place more rapidly than in others. of course the progress of events in any experiment necessarily depends on various causes, such as the size of the containing vessels, but it is easy to see that when experiments with different substances are carried out under similar conditions, however these "similar conditions" be defined, the rates of diffusion must be capable of numerical comparison, and the results must be expressible in terms of at least one physical quantity, which for any two substances can be called their coefficient of diffusion. how to select this quantity we shall see later. 2 _quantitative methods of observing diffusion._--the simplest plan of determining the progress of diffusion between two liquids would be to draw off and examine portions from different strata at some stage in the process; the disturbance produced would, however, interfere with the subsequent process of diffusion, and the observations could not be continued. by placing in the liquid column hollow glass beads of different average densities, and observing at what height they remain suspended, it is possible to trace the variations of density of the liquid column at different depths, and different times. in this method, which was originally introduced by lord kelvin, difficulties were caused by the adherence of small air bubbles to the beads. in general, optical methods are the most capable of giving exact results, and the following may be distinguished, (a) _by refraction in a horizontal plane._ if the containing vessel is in the form of a prism, the deviation of a horizontal ray of light in passing through the prism determines the index of refraction, and consequently the density of the stratum through which the ray passes, (b) _by refraction in a vertical plane._ owing to the density varying with the depth, a horizontal ray entering the liquid also undergoes a small vertical deviation, being bent downwards towards the layers of greater density. the observation of this vertical deviation determines not the actual density, but its rate of variation with the depth, i.e. the "density gradient" at any point, (c) _by the saccharimeter._ in the cases of solutions of sugar, which cause rotation of the plane of polarized light, the density of the sugar at any depth may be determined by observing the corresponding angle of rotation, this was done originally by w. voigt. 3. _elementary definitions of coefficient of diffusion._--the simplest case of diffusion is that of a substance, say a gas, diffusing in the interior of a homogeneous solid medium, which remains at rest, when no external forces act on the system. we may regard it as the result of experience that: (1) if the density of the diffusing substance is everywhere the same no diffusion takes place, and (2) if the density of the diffusing substance is different at different points, diffusion will take place from places of greater to those of lesser density, and will not cease until the density is everywhere the same. it follows that the rate of flow of the diffusing substance at any point in any direction must depend on the density gradient at that point in that direction, i.e. on the rate at which the density of the diffusing substance decreases as we move in that direction. we may define the _coefficient of diffusion_ as the ratio of the total mass per unit area which flows across any small section, to the rate of decrease of the density per unit distance in a direction perpendicular to that section. in the case of steady diffusion parallel to the axis of x, if [rho] be the density of the diffusing substance, and q the mass which flows across a unit of area in a plane perpendicular to the axis of x, then the density gradient is -d[rho]/dx and the ratio of q to this is called the "coefficient of diffusion." by what has been said this ratio remains finite, however small the actual gradient and flow may be., and it is natural to assume, at any rate as a first approximation, that it is constant as far as the quantities in question are concerned. thus if the coefficient of diffusion be denoted by k we have q= -k(d[rho]/dx). further, the rate at which the quantity of substance is increasing in an element between the distances x and x+dx is equal to the difference of the rates of flow in and out of the two faces, whence as in hydrodynamics, we have d[rho]/dt =-dq/dx. it follows that the equation of diffusion in this case assumes the form d[rho] d / d[rho] \ ------ = -- ( k ------ ), dt dx \ dx / which is identical with the equations representing conduction of heat, flow of electricity and other physical phenomena. for motion in three dimensions we have in like manner d[rho] d / d[rho]\ d / d[rho]\ d / d[rho]\ ------ = -- ( k ------ ) + -- ( k ------ ) + -- ( k ------ ); dt dx \ dx / dy \ dy / dz \ dz / and the corresponding equations in electricity and heat for anisotropic substances would be available to account for any parallel phenomena, which may arise, or might be conceived, to exist in connexion with diffusion through a crystalline solid. in the case of a very dilute solution, the coefficient of diffusion of the dissolved substance can be defined in the same way as when the diffusion takes place in a solid, because the effects of diffusion will not have any perceptible influence on the solvent, and the latter may therefore be regarded as remaining practically at rest. but in most cases of diffusion between two fluids, both of the fluids are in motion, and hence there is far greater difficulty in determining the motion, and even in defining the coefficient of diffusion. it is important to notice in the first instance, that it is only the relative motion of the two substances which constitutes diffusion. thus when a current of air is blowing, under ordinary circumstances the changes which take place are purely mechanical, and do not depend on the separate diffusions of the oxygen and nitrogen of which the air is mainly composed. it is only when two gases are flowing with unequal velocity, that is, when they have a relative motion, that these changes of relative distribution, which are called diffusion, take place. the best way out of the difficulty is to investigate the separate motions of the two fluids, taking account of the mechanical actions exerted on them, and supposing that the mutual action of the fluids causes either fluid to resist the relative motion of the other. 4. _the coefficient of resistance._--let us call the two diffusing fluids a and b. if b were absent, the motion of the fluid a would be determined entirely by the variations of pressure of the fluid a, and by the external forces, such as that due to gravity acting on a. similarly if a were absent, the motion of b would be determined entirely by the variations of pressure due to the fluid b, and by the external forces acting on b. when both fluids are mixed together, each fluid tends to resist the relative motion of the other, and by the law of equality of action and reaction, the resistance which a experiences from b is everywhere equal and opposite to the resistance which b experiences from a. if the amount of this resistance per unit volume be divided by the relative velocity of the two fluids, and also by the product of their densities, the quotient is called the "coefficient of resistance." if then [rho]1, [rho]2 are the densities cf the two fluids, u1, u2 their velocities, c the coefficient of resistance, then the portion of the fluid a contained in a small element of volume v will experience from the fluid b a resistance c[rho]1[rho]2v(u1- u2), and the fluid b contained in the same volume element will experience from the fluid a an equal and opposite resistance, c[rho]1[rho]2v(u2 - u1). this definition implies the following laws of resistance to diffusion, which must be regarded as based on experience, and not as self-evident truths: (1) each fluid tends to assume, so far as diffusion is concerned, the same equuibrium distribution that it would assume if its motion were unresisted by the presence of the other fluid. (of course, the mutual attraction of gravitation of the two fluids might affect the final distribution, but this is practically negligible. leaving such actions as this out of account the following statement is correct.) in a state of equilibrium, the density of each fluid at any point thus depends only on the partial pressure of that fluid alone, and is the same as if the other fluids were absent. it does not depend on the partial pressures of the other fluids. if this were not the case, the resistance to diffusion would be analogous to friction, and would contain terms which were independent of the relative velocity u2 - u1. (2) for slow motions the resistance to diffusion is (approximately at any rate) proportional to the relative velocity. (3) the coefficient of resistance c is not necessarily always constant; it may, for example, and, in general, does, depend on the temperature. if we form the equations of hydrodynamics for the different fluids occurring in any mixture, taking account of diffusion, but neglecting viscosity, and using suffixes 1, 2 to denote the separate fluids, these assume the form given by james clerk maxwell ("diffusion," in _ency. brit._, 9th ed.):-- du1 dp1 [rho] --- + --- - x1[rho]1 + c12[rho]1[rho]2(u1 - u2) + &c. = 0, dt dx where du1 du1 du1 du1 du1 --- = --- + u1 --- + v1 --- + w1 ---, dt dt dx dy dz and these equations imply that when diffusion and other motions cease, the fluids satisfy the separate conditions of equilibrium dp1/dx - x1[rho]1 = 0. the assumption made in the following account is that terms such as du1/dt may be neglected in the cases considered. a further property based on experience is that the motions set up in a mixture by diffusion are very slow compared with those set up by mechanical actions, such as differences of pressure. thus, if two gases at equal temperature and pressure be allowed to mix by diffusion, the heavier gas being below the lighter, the process will take a long time; on the other hand, if two gases, or parts of the same gas, at different pressures be connected, equalization of pressure will take place almost immediately. it follows from this property that the forces required to overcome the "inertia" of the fluids in the motions due to diffusion are quite imperceptible. at any stage of the process, therefore, any one of the diffusing fluids may be regarded as in equilibrium under the action of its own partial pressure, the external forces to which it is subjected and the resistance to diffusion of the other fluids. 5. _slow diffusion of two gases. relation between the coefficients of resistance and of diffusion._--we now suppose the diffusing substances to be two gases which obey boyle's law, and that diffusion takes place in a closed cylinder or tube of unit sectional area at constant temperature, the surfaces of equal density being perpendicular to the axis of the cylinder, so that the direction of diffusion is along the length of the cylinder, and we suppose no external forces, such as gravity, to act on the system. the densities of the gases are denoted by [rho]1, [rho]2, their velocities of diffusion by u1, u2, and if their partial pressures are p1, p2, we have by boyle's law p1 = k1[rho]1, p2 = k2[rho]2, where k1, k2 are constants for the two gases, the temperature being constant. the axis of the cylinder is taken as the axis of x. from the considerations of the preceding section, the effects of inertia of the diffusing gases may be neglected, and at any instant of the process either of the gases is to be treated as kept in equilibrium by its partial pressure and the resistance to diffusion produced by the other gas. calling this resistance per unit volume r, and putting r = c[rho]1[rho]2(u1 - u2), where c is the coefficient of resistance, the equations of equilibrium give dp1 dp2 --- + c[rho]1[rho]2(u1 - u2)= 0, and --- + c[rho]1[rho]2(u2 - u1)= 0 (1). dx dx these involve dp1 dp2 --- + --- = 0 or p1 + p2 = p (2) dx dx where p is the total pressure of the mixture, and is everywhere constant, consistently with the conditions of mechanical equilibrium. now dp1/dx is the pressure-gradient of the first gas, and is, by boyle's law, equal to k1 times the corresponding density-gradient. again [rho]1u1 is the mass of gas flowing across any section per unit time, and k1[rho]1u1 or p1u1 can be regarded as representing the flux of partial pressure produced by the motion of the gas. since the total pressure is everywhere constant, and the ends of the cylinder are supposed fixed, the fluxes of partial pressure due to the two gases are equal and opposite, so that p1u1 + p2u2 = 0 or k1[rho]1u1 + k2[rho]2u2 = 0 (3). from (2) (3) we find by elementary algebra u1/p2 = - u2/p1 = (u1 - u2)/(p1 + p2) = (u1 - u2)/p, and therefore p2u1 = - p2u2 = p1p2(u1 - u2)/p = k1k2[rho]1[rho]2(u1 - u2)/p hence equations (1) (2) gives dp1 cp dp2 cp --- + ---- (p1u1) = 0, and --- + ---- (p2u2) = 0; dx k1k2 dx k1k2 whence also substituting p1 = k1[rho]1, p2 = k2[rho]2, and by transposing k1k2 d[rho]1 k1k2 d[rho]2 [rho]1u1 = - ---- -------, and [rho]2u2 = - ---- -------. cp dx cp dx we may now define the "coefficient of diffusion" of either gas as the ratio of the rate of flow of that gas to its density-gradient. with this definition, the coefficients of diffusion of both the gases in a mixture are equal, each being equal to k1k2/cp. the ratios of the fluxes of partial pressure to the corresponding pressure-gradients are also equal to the same coefficient. calling this coefficient k, we also observe that the equations of continuity for the two gases are d[rho]1 d([rho]1u1) d[rho]2 d([rho]2u2) ------- + ----------- = 0, and ------- + ----------- = 0, dt dx dt dx leading to the equations of diffusion d[rho]1 d / d[rho]1\ d[rho]2 d / d[rho]2\ ------- = -- ( k ------- ) , and ------- = -- ( k ------- ), dt dx \ dx / dt dx \ dx / exactly as in the case of diffusion through a solid. if we attempt to treat diffusion in liquids by a similar method, it is, in the first place, necessary to define the "partial pressure" of the components occurring in a liquid mixture. this leads to the conception of "osmotic pressure," which is dealt with in the article solution. for dilute solutions at constant temperature, the assumption that the osmotic pressure is proportional to the density, leads to results agreeing fairly closely with experience, and this fact may be represented by the statement that a substance occurring in a dilute solution behaves like a perfect gas. 6. _relation of the coefficient of diffusion to the units of length and time._--we may write the equation defining k in the form i d[rho] u = -k × ----- ------. [rho] dx here -d[rho]/[rho]dx represents the "percentage rate" at which the density decreases with the distance x; and we thus see that the coefficient of diffusion represents the ratio of the velocity of flow to the percentage rate at which the density decreases with the distance measured in the direction of flow. this percentage rate being of the nature of a number divided by a length, and the velocity being of the nature of a length divided by a time, we may state that k is of two dimensions in length and - 1 in time, i.e. dimensions l2/t. _example 1._ taking k = 0.1423 for carbon dioxide and air (at temperature 0° c. and pressure 76 cm. of mercury) referred to a centimetre and a second as units, we may interpret the result as follows:--supposing in a mixture of carbon dioxide and air, the density of the carbon dioxide decreases by, say, 1, 2 or 3% of itself in a distance of 1 cm., then the corresponding velocities of the diffusing carbon dioxide will be respectively 0.01, 0.02 and 0.03 times 0.1423, that is, 0.001423, 0.002846 and 0.004269 cm. per second in the three cases. _example 2._ if we wished to take a foot and a second as our units, we should have to divide the value of the coefficient of diffusion in example 1 by the square of the number of centimetres in 1 ft., that is, roughly speaking, by 900, giving the new value of k = 0.00016 roughly. 7. _numerical values of the coefficient of diffusion._--the table on p. 258 gives the values of the coefficient of diffusion of several of the principal pairs of gases at a pressure of 76 cm. of mercury, and also of a number of other substances. in the gases the centimetre and second are taken as fundamental units, in other cases the centimetre and day. 8. _irreversible changes accompanying diffusion._--the diffusion of two gases at constant pressure and temperature is a good example of an "irreversible process." the gases always tend to mix, never to separate. in order to separate the gases a change must be effected in the external conditions to which the mixture is subjected, either by liquefying one of the gases, or by separating them by diffusion through a membrane, or by bringing other outside influences to bear on them. in the case of liquids, electrolysis affords a means of separating the constituents of a mixture. every such method involves some change taking place outside the mixture, and this change may be regarded as a "compensating transformation." we thus have an instance of the property that every irreversible change leaves an indelible imprint somewhere or other on the progress of events in the universe. that the process of diffusion obeys the laws of irreversible thermodynamics (if these laws are properly stated) is proved by the fact that the compensating transformations required to separate mixed gases do not essentially involve anything but transformation of energy. the process of allowing gases to mix by diffusion, and then separating them by a compensating transformation, thus constitutes an irreversible cycle, the outside effects of which are that energy somewhere or other must be less capable of transformation than it was before the change. we express this fact by stating that an irreversible process essentially implies a loss of availability. to measure this loss we make use of the laws of thermodynamics, and in particular of lord kelvin's statement that "it is impossible by means of inanimate material agency to derive mechanical effect from any portion of matter by cooling it below the temperature of the coldest of the surrounding objects." +-------------------------------------------+---------+---------------------+--------------+ | substances. | temp. | k. | author. | +-------------------------------------------+---------+---------------------+--------------+ | carbon dioxide and air | 0°c. | 0.1423 cm2/sec. | j. loschmidt.| | " " hydrogen | 0°c. | 0.5558 " | " | | " " oxygen | 0°c. | 0.1409 " | " | | " " carbon monoxide | 0°c. | 0.1406 " | " | | " " marsh gas (methane) | 0°c. | 0.1586 " | " | | " " nitrous oxide | 0°c. | 0.0983 " | " | | hydrogen and oxygen | 0°c. | 0.7214 " | " | | " " carbon monoxide | 0°c. | 0.6422 " | " | | " " sulphur dioxide | 0°c. | 0.4800 " | " | | oxygen and carbon monoxide | 0°c. | 0.1802 " | " | | water and ammonia | 20°c. | 1.250 " | g. hufner. | | " " | 5°c. | 0.822 " | " | | " common salt (density 1.0269) | | 0.355 cm2/hour. | j. graham. | | " " " " |14.33°c. | 1.020, 0.996, 0.972,| " | | | | 0.932 cm2/day. | f. heimbrodt.| | " zinc sulphate (0.312 gm/cm3) | | 0.1162 " | w. seitz. | | " zinc sulphate (normal) | | 0.2355 " | " | | " zinc acetate (double normal) | | 0.1195 " | " | | " zinc formate (half normal) | | 0.4654 " | " | | " cadmium sulphate (double normal)| | 0.2456 " | " | | " glycerin (1/8n, 1⁄2n, 7/8n, 7/8n) |10.14°c. | 0.356, 0.350, 0.342,| f. heimbrodt.| | | | 0.315 cm2/day. | " | | " urea " " |14.83°c. | 0.973, 0.946, 0.926,| " | | | | 0.883 cm2/day. | " | | " hydrochloric acid |14.30°c. | 2.208, 2.331, | " | | | | 2.480 cm2/day | " | | gelatin 20% and ammonia | 17°c. | 127.1 " | a. hagenbach.| | " " carbon dioxide | | 0.845 " | " | | " " nitrous oxide | | 0.509 " | " | | " " oxygen | | 0.230 " | " | | " " hydrogen | | 0.0565 " | " | +-------------------------------------------+---------+---------------------+--------------+ let us now assume that we have any syste m such as the gases above considered, and that it is in the presence of an indefinitely extended medium which we shall call the "auxiliary medium." if heat be taken from any part of the system, only part of this heat can be converted into work by means of thermodynamic engines; and the rest will be given to the auxiliary medium, and will constitute unavailable energy or waste. to understand what this means, we may consider the case of a condensing steam engine. only part of the energy liberated by the combustion of the coal is available for driving the engine, the rest takes the form of heat imparted to the condenser. the colder the condenser the more efficient is the engine, and the smaller is the quantity of waste. the amount of unavailable energy associated with any given transformation is proportional to the absolute temperature of the auxiliary medium. when divided by that temperature the quotient is called the change of "entropy" associated with the given change (see thermodynamics). thus if a body at temperature t receives a quantity of heat q, and if t0 is the temperature of the auxiliary medium, the quantity of work which could be obtained from q by means of ideal thermodynamic engines would be q(1 - t0/t), and the balance, which is qt0/t, would take the form of unavailable or waste energy given to the medium. the quotient of this, when divided by t0, is q/t, and this represents the quantity of entropy associated with q units of heat at temperature t. any irreversible change for which a compensating transformation of energy exists represents, therefore, an increase of unavailable energy, which is measurable in terms of entropy. the increase of entropy is independent of the temperature of the auxiliary medium. it thus affords a measure of the extent to which energy has run to waste during the change. moreover, when a body is heated, the increase of entropy is the factor which determines how much of the energy imparted to the body is unavailable for conversion into work under given conditions. in all cases we have increase of unavailable energy ------------------------------- = increase of entropy. temperature of auxiliary medium when diffusion takes place between two gases inside a closed vessel at uniform pressure and temperature no energy in the form of heat or work is received from without, and hence the entropy gained by the gases from without is zero. but the irreversible processes inside the vessel may involve a gain of entropy, and this can only be estimated by examining by what means mixed gases can be separated, and, in particular, under what conditions the process of mixing and separating the gases could (theoretically) be made reversible. 9. _evidence derived from liquefaction of one or both of the gases._--the gases in a mixture can often be separated by liquefying, or even solidifying, one or both of the components. in connexion with this property we have the important law according to which "the pressure of a vapour in equilibrium with its liquid depends only on the temperature and is independent of the pressures of any other gases or vapours which may be mixed with it." thus if two closed vessels be taken containing some water and one be exhausted, the other containing air, and if the temperatures be equal, evaporation will go on until the pressure of the vapour in the exhausted vessel is equal to its _partial_ pressure in the other vessel, notwithstanding the fact that the _total_ pressure in the latter vessel is greater by the pressure of the air. to separate mixed gases by liquefaction, they must be compressed and cooled till one separates in the form of a liquid. if no changes are to take place outside the system, the separate components must be allowed to expand until the work of expansion is equal to the work of compression, and the heat given out in compression is reabsorbed in expansion. the process may be made as nearly reversible as we like by performing the operations so slowly that the substances are practically in a state of equilibrium at every stage. this is a consequence of an important axiom in thermodynamics according to which "any small change in the neighbourhood of a state of equilibrium is to a first approximation reversible." suppose now that at any stage of the compression the partial pressures of the two gases are p1 and p2, and that the volume is changed from v to v - dv. the work of compression is (p1 + p2)dv, and this work will be restored at the corresponding stage if each of the separated gases increases in volume from v - dv to v. the ultimate state of the separated gases will thus be one in which each gas occupies the volume v originally occupied by the mixture. we may now obtain an estimate of the amount of energy rendered unavailable by diffusion. we suppose two gases occupying volumes v1 and v2 at equal pressure p to mix by diffusion, so that the final volume is v1 + v2. then if before mixing each gas had been allowed to expand till its volume was v1 + v2, work would have been done in the expansion, and the gases could still have been mixed by a reversal of the process above described. in the actual diffusion this work of expansion is lost, and represents energy rendered unavailable at the temperature at which diffusion takes place. when divided by that temperature the quotient gives the increase of entropy. thus the irreversible processes, and, in particular, the entropy changes associated with diffusion of two gases at uniform pressure, are the same as would take place if each of the gases in turn were to expand by rushing into a vacuum, till it occupied the whole volume of the mixture. a more rigorous proof involves considerations of the thermodynamic potentials, following the methods of j. willard gibbs (see energetics). another way in which two or more mixed gases can be separated is by placing them in the presence of a liquid which can freely absorb one of the gases, but in which the other gas or gases are insoluble. here again it is found by experience that when equilibrium exists at a given temperature between the dissolved and undissolved portions of the first gas, the partial pressure of that gas in the mixture depends on the temperature alone, and is independent of the partial pressures of the insoluble gases with which it is mixed, so that the conclusions are the same as before. 10. _diffusion through a membrane or partition. theory of the semi-permeable membrane._--it has been pointed out that diffusion of gases frequently takes place in the interior of solids; moreover, different gases behave differently with respect to the same solid at the same temperature. a membrane or partition formed of such a solid can therefore be used to effect a more or less complete separation of gases from a mixture. this method is employed commercially for extracting oxygen from the atmosphere, in particular for use in projection lanterns where a high degree of purity is not required. a similar method is often applied to liquids and solutions and is known as "dialysis." in such cases as can be tested experimentally it has been found that a gas always tends to pass through a membrane from the side where its density, and therefore its partial pressure, is greater to the side where it is less; so that for equilibrium the partial pressures on the two sides must be equal. this result is unaffected by the presence of other gases on one or both sides of the membrane. for example, if different gases at the same pressure are separated by a partition through which one gas can pass more rapidly than the other, the diffusion will give rise to a difference of pressure on the two sides, which is capable of doing mechanical work in moving the partition. in evidence of this conclusion max planck quotes a test experiment made by him in the physical institute of the university of munich in 1883, depending on the fact that platinum foil at white heat is permeable to hydrogen but impermeable to air, so that if a platinum tube filled with hydrogen be heated the hydrogen will diffuse out, leaving a vacuum. the details of the experiment may be quoted here:--"a glass tube of about 5 mm. internal diameter, blown out to a bulb at the middle, was provided with a stop-cock at one end. to the other a platinum tube 10 cm. long was fastened, and closed at the end. the whole tube was exhausted by a mercury pump, filled with hydrogen at ordinary atmospheric pressure, and then closed. the closed end of the platinum portion was then heated in a horizontal position by a bunsen burner. the connexion between the glass and platinum tubes, having been made by means of sealing-wax, had to be kept cool by a continuous current of water to prevent the softening of the wax. after four hours the tube was taken from the flame, cooled to the temperature of the room, and the stop-cock opened under mercury. the mercury rose rapidly, almost completely filling the tube, proving that the tube had been very nearly exhausted." [illustration] in order that diffusion through a membrane may be reversible so far as a particular gas is concerned, the process must take place so slowly that equilibrium is set up at every stage (see § 9 above). in order to separate one gas from another consistently with this condition it is necessary that no diffusion of the latter gas should accompany the process. the name "semi-permeable" is applied to an ideal membrane or partition through which one gas can pass, and which offers an insuperable barrier to any diffusion whatever of a second gas. by means of two semi-permeable partitions acting oppositely with respect to two different gases a and b these gases could be mixed or separated by reversible methods. the annexed figure shows a diagrammatic representation of the process. we suppose the gases contained in a cylindrical tube; p, q, r, s are four pistons, of which p and r are joined to one connecting rod, q and s to another. p, s are impermeable to both gases; q is semi-permeable, allowing the gas a to pass through but not b, similarly r allows the gas b to pass through but not a. the distance pr is equal to the distance qs, so that if the rods are pushed towards each other as far as they will go, p and q will be in contact, as also r and s. imagine the space rq filled with a mixture of the two gases under these conditions. then by slowly drawing the connecting rods apart until r, q touch, the gas a will pass into the space pq, and b will pass into the space rs, and the gases will finally be completely separated; similarly, by pushing the connecting rods together, the two gases will be remixed in the space rq. by performing the operations slowly enough we may make the processes as nearly reversible as we please, so that no available energy is lost in either change. the gas a being at every instant in equilibrium on the two sides of the piston q, its density, and therefore its partial pressure, is the same on both sides, and the same is true regarding the gas b on the two sides of r. also _no work is done in moving the pistons_, for the partial pressures of b on the two sides of r balance each other, consequently, the resultant thrust on r is due to the gas a alone, and is equal and opposite to its resultant thrust on p, so that the connecting rods are at every instant in a state of mechanical equilibrium so far as the pressures of the gases a and b are concerned. we conclude that in the reversible separation of the gases by this method at constant temperature without the production or absorption of mechanical work, the densities and the partial pressures of the two separated gases are the same as they were in the mixture. these conclusions are in entire agreement with those of the preceding section. if this agreement did not exist it would be possible, theoretically, to obtain perpetual motion from the gases in a way that would be inconsistent with the second law of thermodynamics. most physicists admit, as planck does, that it is impossible to obtain an ideal semi-permeable substance; indeed such a substance would necessarily have to possess an infinitely great resistance to diffusion for such gases as could not penetrate it. but in an experiment performed under actual conditions the losses of available energy arising from this cause would be attributable to the imperfect efficiency of the partitions and not to the gases themselves; moreover, these losses are, in every case, found to be completely in accordance with the laws of irreversible thermodynamics. the reasoning in this article being somewhat condensed the reader must necessarily be referred to treatises on thermodynamics for further information on points of detail connected with the argument. even when he consults these treatises he may find some points omitted which have been examined in full detail at some time or other, but are not sufficiently often raised to require mention in print. ii. _kinetic models of diffusion._--imagine in the first instance that a very large number of red balls are distributed over one half of a billiard table, and an equal number of white balls over the other half. if the balls are set in motion with different velocities in various directions, diffusion will take place, the red balls finding their way among the white ones, and vice versa; and the process will be retarded by collisions between the balls. the simplest model of a perfect gas studied in the kinetic theory of gases (see molecule) differs from the above illustration in that the bodies representing the molecules move in space instead of in a plane, and, unlike billiard balls, their motion is unresisted, and they are perfectly elastic, so that no kinetic energy is lost either during their free motions, or at a collision. the mathematical analysis connected with the application of the kinetic theory to diffusion is very long and cumbersome. we shall therefore confine our attention to regarding a medium formed of elastic spheres as a mechanical model, by which the most important features of diffusion can be illustrated. we shall assume the results of the kinetic theory, according to which:--(1) in a dynamical model of a perfect gas the mean kinetic energy of translation of the molecules represents the absolute temperature of the gas. (2) the pressure at any point is proportional to the product of the number of molecules in unit volume about that point into the mean square of the velocity. (the mean square of the velocity is different from but proportional to the square of the mean velocity, and in the subsequent arguments either of these two quantities can generally be taken.) (3) in a gas mixture represented by a mixture of molecules of unequal masses, the mean kinetic energies of the different kinds are equal. consider now the problem of diffusion in a region containing two kinds of molecules a and b of unequal mass. the molecules of a in the neighbourhood of any point will, by their motion, spread out in every direction until they come into collision with other molecules of either kind, and this spreading out from every point of the medium will give rise to diffusion. if we imagine the velocities of the a molecules to be equally distributed in all directions, as they would be in a homogeneous mixture, it is obvious that the process of diffusion will be greater, _ceteris paribus_, the greater the velocity of the molecules, and the greater the length of the free path before a collision takes place. if we assume consistently with this, that the coefficient of diffusion of the gas a is proportional to the mean value of wala, where wa is the velocity and la is the length of the path of a molecule of a, this expression for the coefficient of diffusion is of the right dimensions in length and time. if, moreover, we observe that when diffusion takes place in a fixed direction, say that of the axis of x, it depends only on the resolved part of the velocity and length of path in that direction: this hypothesis readily leads to our taking the mean value of 1/3w_a l_a as the coefficient of diffusion for the gas a. this value was obtained by o. e. meyer and others. unfortunately, however, it makes the coefficients of diffusion unequal for the two gases, a result inconsistent with that obtained above from considerations of the coefficient of resistance, and leading to the consequence that differences of pressure would be set up in different parts of the gas. to equalize these differences of pressure, meyer assumed that a counter current is set up, this current being, of course, very slow in practice; and j. stefan assumed that the diffusion of one gas was not affected by collisions between molecules of the _same gas_. when the molecules are mixed in equal proportions both hypotheses lead to the value 1/6([w_a l_a] + [w_b l_b]), (square brackets denoting mean values). when one gas preponderates largely over the other, the phenomena of diffusion are too difficult of observation to allow of accurate experimental tests being made. moreover, in this case no difference exists unless the molecules are different in size or mass. instead of supposing a velocity of translation added after the mathematical calculations have been performed, a better plan is to assume from the outset that the molecules of the two gases have small velocities of translation in opposite directions, superposed on the distribution of velocity, which would occur in a medium representing a gas at rest. when a collision occurs between molecules of different gases a transference of momentum takes place between them, and the quantity of momentum so transferred in one second in a unit of volume gives a dynamical measure of the resistance to diffusion. it is to be observed that, however small the relative velocity of the gases a and b, it plays an all-important part in determining the coefficient of resistance; for without such relative motion, and with the velocities evenly distributed in all directions, no transference of momentum could take place. the coefficient of resistance being found, the motion of each of the two gases may be discussed separately. one of the most important consequences of the kinetic theory is that if the volume be kept constant the coefficient of diffusion varies as the square root of the absolute temperature. to prove this, we merely have to imagine the velocity of each molecule to be suddenly increased n fold; the subsequent processes, including diffusion, will then go on n times as fast; and the temperature t, being proportional to the kinetic energy, and therefore to the square of the velocity, will be increased n2 fold. thus k, the coefficient of diffusion, varies as [sqrt]t. the relation of k to the density when the temperature remains constant is more difficult to discuss, but it may be sufficient to notice that if the number of molecules is increased n fold, the chances of a collision are n times as great, and the distance traversed between collisions is (not _therefore_ but as the result of more detailed reasoning) on the average 1/n of what it was before. thus the free path, and therefore the coefficient of diffusion, varies inversely as the density, or directly as the volume. if the pressure p and temperature t be taken as variables, k varies inversely as p and directly as [sqrt]t3. now according to the experiments first made by j. c. maxwell and j. loschmidt, it appeared that with constant density k was proportional to t more nearly than to [sqrt]t. the inference is that in this respect a medium formed of colliding spheres fails to give a correct mechanical model of gases. it has been found by l. boltzmann, maxwell and others that a system of particles whose mutual actions vary according to the inverse fifth power of the distance between them represents more correctly the relation between the coefficient of diffusion and temperature in actual gases. other recent theories of diffusion have been advanced by m. thiesen, p. langevin and w. sutherland. on the other hand, j. thovert finds experimental evidence that the coefficient of diffusion is proportional to molecular velocity in the cases examined of non-electrolytes dissolved in water at 18° at 2.5 grams per litre. bibliography.--the best introduction to the study of theories of diffusion is afforded by o. e. meyer's kinetic _theory of gases_, translated by robert e. baynes (london, 1899). the mathematical portion, though sufficient for ordinary purposes, is mostly of the simplest possible character. another useful treatise is r. ruhlmann's _handbuch der mechanischen warmetheorie_ (brunswick, 1885). for a shorter sketch the reader may refer to j. c. maxwell's _theory of heat_, chaps, xix. and xxii., or numerous other treatises on physics. the theory of the semi-permeable membrane is discussed by m. planck in his _treatise on thermodynamics_, english translation by a. ogg (1903), also in treatises on thermodynamics by w. voigt and other writers. for a more detailed study of diffusion in general the following papers may be consulted:--l. boltzmann, "zur integration der diffusionsgleichung," _sitzung. der k. bayer. akad math.-phys. klasse_ (may 1894); t. des coudres, "diffusionsvorgange in einem zylinder," _wied. ann._ lv. (1895), p. 213; j. loschmidt, "experimentaluntersuchungen uber diffusion," _wien. sitz._ lxi., lxii. (1870); j. stefan, "gleichgewicht und ... diffusion von gasmengen," _wien. sitz._ lxiii., "dynamische theorie der diffusion," _wien. sitz._ lxv. (april 1872); m. toepler, "gas-diffusion," _wied. ann._ lviii. (1896), p. 599; a. wretschko, "experimentaluntersuchungen uber die diffusion von gasmengen," _wien. sitz._ lxii. the mathematical theory of diffusion, according to the kinetic theory of gases, has been treated by a number of different methods, and for the study of these the reader may consult l. boltzmann, _vorlesungen uber gastheorie_ (leipzig, 1896-1898); s. h. burbury, _kinetic theory of gases_ (cambridge, 1899), and papers by l. boltzmann in _wien. sitz._ lxxxvi. (1882), lxxxvii. (1883); p. g. tait, "foundations of the kinetic theory of gases," _trans. r.s.e._ xxxiii., xxxv., xxvi., or _scientific papers_, ii. (cambridge, 1900). for recent work reference should be made to the current issues of _science abstracts_ (london), and entries under the heading "diffusion" will be found in the general index at the end of each volume. (g. h. br.)