Вы находитесь на странице: 1из 8

Finite state machine

o Finite state machine is used to recognize patterns.


o Finite automata machine takes the string of symbol as input and changes its state
accordingly. In the input, when a desired symbol is found then the transition occurs.
o While transition, the automata can either move to the next state or stay in the same
state.
o FA has two states: accept state or reject state. When the input string is successfully
processed and the automata reached its final state then it will accept.

A finite automata consists of following:

Q: finite set of states


∑: finite set of input symbol
q0: initial state
F: final state
δ: Transition function

Transition function can be define as δ: Q x ∑ →Q

FA is characterized into two ways:

1. DFA (finite automata)


2. NDFA (non deterministic finite automata)

DFA:-

DFA stands for Deterministic Finite Automata. Deterministic refers to the uniqueness of the
computation. In DFA, the input character goes to one state only. DFA doesn't accept the null
move that means the DFA cannot change state without any input character.

DFA has five tuples {Q, ∑, q0, F, δ}

Q: set of all states


∑: finite set of input symbol where δ: Q x ∑ →Q
q0: initial state
F: final state
δ: Transition function

NDFA:-

NDFA refer to the Non Deterministic Finite Automata. It is used to transit the any number of
states for a particular input. NDFA accepts the NULL move that means it can change state
without reading the symbols.

NDFA also has five states same as DFA. But NDFA has different transition function.
Transition function of NDFA can be defined as δ: Q x ∑ →2Q
Regular expression

o Regular expression is a sequence of pattern that defines a string. It is used to denote


regular languages.
o It is also used to match character combinations in strings. String searching algorithm
used this pattern to find the operations on string.
o In regular expression, x* means zero or more occurrence of x. It can generate {e, x,
xx, xxx, xxxx,.....}
o In regular expression, x+ means one or more occurrence of x. It can generate {x,
xx, xxx, xxxx,.....}

Operations on Regular Language

The various operations on regular language are:

Union: If L and M are two regular languages then their union L U M is also a union.

L U M = {s | s is in L or s is in M}

Intersection: If L and M are two regular languages then their intersection is also an
intersection.

L ⋂ M = {st | s is in L and t is in M}

Kleene closure: If L is a regular language then its kleene closure L1* will also be a regular
language.

L* = Zero or more occurrence of language L.

Example Write the regular expression for the language:

L = {abn w:n ≥ 3, w ∈ (a,b)+}

Solution: The string of language L starts with "a" followed by atleast three b's. Itcontains
atleast one "a" or one "b" that is string are like abbba, abbbbbba, abbbbbbbb, abbbb.....a

So regular expression is: r= ab3b* (a+b)+

Here + is a positive closure i.e. (a+b)+ = (a+b)* - ∈

Context Free grammar :-

Context free grammar is a formal grammar which is used to generate all possible strings in
a given formal language.
Context free grammar G can be defined by four tuples as:

G= (V, T, P, S)

Where,

G describes the grammar

T describes a finite set of terminal symbols.

V describes a finite set of non-terminal symbols

P describes a set of production rules

S is the start symbol.

In CFG, the start symbol is used to derive the string. You can derive the string by
repeatedly replacing a non-terminal by the right hand side of the production, until all non-
terminal have been replaced by terminal symbols.

Example:

L= {wcwR | w € (a, b)*}

Production rules:

1. S → aSa
2. S → bSb
3. S → c

Now check that abbcbba string can be derived from the given CFG.

1. S ⇒ aSa
2. S ⇒ abSba
3. S ⇒ abbSbba
4. S ⇒ abbcbba

By applying the production S → aSa, S → bSb recursively and finally applying the
production S → c, we get the string abbcbba

Capabilities of CFG:-

There are the various capabilities of CFG:

o Context free grammar is useful to describe most of the programming languages.


o If the grammar is properly designed then an efficientparser can be constructed
automatically.
o Using the features of associatively & precedence information, suitable grammars for
expressions can be constructed.
o Context free grammar is capable of describing nested structures like: balanced
parentheses, matching begin-end, corresponding if-then-else's & so on.

Derivation

Derivation is a sequence of production rules. It is used to get the input string through these
production rules. During parsing we have to take two decisions. These are as follows:

o We have to decide the non-terminal which is to be replaced.


o We have to decide the production rule by which the non-terminal will be replaced.

We have two options to decide which non-terminal to be replaced with production rule.

Left-most Derivation

In the left most derivation, the input is scanned and replaced with the production rule from
left to right. So in left most derivatives we read the input string from left to right.

Example:

Production rules:

1. S  S + S
2. S  S - S
3. S  a | b |c

Input:

a - b + c

The left-most derivation is:

1. S  S+S
2. S  S-S+S
3. S  a-S+S
4. S  a-b+S
5. S  a-b+c

Right-most Derivation

In the right most derivation, the input is scanned and replaced with the production rule from
right to left. So in right most derivatives we read the input string from right to left.

Example:
1. S  S + S
2. S  S - S
3. Sa | b |c

Input: a - b + c

The right-most derivation is:

1. S  S-S
2. S  S-S+S
3. S  S-S+c
4. S  S-b+c
5. S  a-b+c

Parse tree:-

o Parse tree is the graphical representation of symbol. The symbol can


be terminal or non-terminal.
o In parsing, the string is derived using the start symbol. The root of the
parse tree is that start symbol.
o It is the graphical representation of symbol that can be terminals or
non-terminals.
o Parse tree follows the precedence of operators. The deepest sub-tree
traversed first. So, the operator in the parent node has less
precedence over the operator in the sub-tree.

The parse tree follows these points:


o All leaf nodes have to be terminals.
o All interior nodes have to be non-terminals.
o In-order traversal gives original input string.

Example: Production rules: T T + T | T * T

T  a|b|c

Input: a * b + c

Step 1:
Step 2:

Step 3:

Step 4:

Step 5:

Ambiguity:-
A grammar is said to be ambiguous if there exists more than one leftmost derivation or
more than one rightmost derivative or more than one parse tree for the given input string.
If the grammar is not ambiguous then it is called unambiguous.

Example:

S  aSb | SS

S∈

For the string aabb, the above grammar generates two parse trees:

If the grammar has ambiguity then it is not good for a compiler construction. No method
can automatically detect and remove the ambiguity but you can remove ambiguity by re-
writing the whole grammar without ambiguity.

Elimination of Left Recursion:

A grammar is left recursion if it has a non-terminal A such that there is a derivation A → Aa/b for some
string a.

. A left-recursive pair of productions A → Aa/b could be replaced by the non-left-recursive productions;-

A→ bA'

A'→ αA'| ε

A→Aα1| Aα2|..........Aαm |β1| β2|........ |βn can be re-written as:-

A→ β1A'| β2A'|........ | βnA' A'→α1A'| α2A'|.......... αmA' | ε

Left Factoring:-

If A αβ1|αβ2 are two A-productions and the input begins with a non empty string derived from α, we
do not know whether to expand A to αβ1 or to αβ2. We may defer decision by expanding A to αA'. Then
after seeing the input derived from a we expand A' to β1or to β2
AαA'

A'β1| β2

In a more standard way Aαβ1| αβ2|..............| αβn|γ , may be left factored as:-

AαA'|γ

A' β1| β2.............|βn

Вам также может понравиться